Instructions to use srpone/zoowork-shopranker-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use srpone/zoowork-shopranker-4b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("srpone/zoowork-shopranker-4b") model = PeftModel.from_pretrained(base_model, "srpone/zoowork-shopranker-4b") - Notebooks
- Google Colab
- Kaggle
ZooWork-ShopRanker-4B
An e-commerce reranker aligned to judged shopper preference rather than topical
relevance, from the paper ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce
Reranker. This repository is self-contained: it holds the
unmodified Qwen/Qwen3-Reranker-4B weights together with our LoRA
adapter, so everything loads from this one repo id -- no separate base download. Same scoring
interface as the base: one forward pass per (query, product), works at any k, and scores
are cacheable per pair.
🚀 An optimized commercial version of the ZooWork-ShopRanker rerankers is available through the zoodata.ai API.
Results on ShopRank-Bench
Pairwise accuracy in percent on ShopRank-Bench (10,511 judged preference pairs from 2,991 queries); intervals are 95% query-clustered bootstrap CIs. Tiers are the number of judge families (of three) that committed to the label.
| Model | Product text | Overall [95% CI] | Gold (3/3) | Silver (2/3) | Bronze (1/3) |
|---|---|---|---|---|---|
| ZooWork-ShopRanker-4B | Structured | 81.2 [80.3, 82.0] | 96.3 | 84.3 | 71.4 |
Qwen3-Reranker-4B (base) |
Structured | 77.5 [76.5, 78.5] | 92.7 | 81.4 | 66.7 |
| ZooWork-ShopRanker-4B | Natural language | 79.5 [78.6, 80.4] | 95.4 | 82.1 | 69.8 |
Qwen3-Reranker-4B (base) |
Natural language | 77.9 [76.9, 78.9] | 92.1 | 81.0 | 68.5 |
Significantly beats the strongest open baseline (Jina-reranker-m0: 79.2 structured, 78.1 natural language) in both formats and its own base by 3.7 points on structured text.
Report per tier: every reranker loses 20–25 points from gold to bronze, so an aggregate score is dominated by pairs nothing gets wrong.
Serving cost
One NVIDIA H200, fp16, 1024-token limit, 256 real (query, product) pairs, adapter merged into the weights (LoRA adds no inference cost once merged).
| Params | Batch-1 p50 | Batch-1 p95 | Docs/s (batch 32) | Peak memory |
|---|---|---|---|---|
| 4.02B | 25.2 ms | 33.3 ms | 154.5 | 11.2 GB |
Usage
Requires pip install transformers peft. Load the base weights and the adapter from this
repo as shown: wrapping with PeftModel keeps the adapter in fp32, which reproduces the paper
numbers exactly (loading the repo id with from_pretrained alone also attaches the adapter,
but in bf16, and drifts by up to ~0.3 points).
The score is P(yes) = softmax over the bare no / yes logits at the last position of
the native Qwen3-Reranker prompt -- exactly the scorer behind the numbers above.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "srpone/zoowork-shopranker-4b" # base weights + LoRA adapter, both in this repo
tok = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)
# Re-wrapping keeps the adapter in fp32, as in the paper. peft may warn "Already found a
# `peft_config` attribute"; that is expected here -- exactly one adapter ends up loaded.
model = PeftModel.from_pretrained(model, model_id).cuda().eval()
PREFIX = ("<|im_start|>system\nJudge whether the Document meets the requirements based on "
"the Query and the Instruct provided. Note that the answer can only be \"yes\" "
"or \"no\".<|im_end|>\n<|im_start|>user\n")
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCTION = "Given a web search query, retrieve relevant passages that answer the query"
YES, NO = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")
@torch.no_grad()
def score(query: str, products: list[str]) -> list[float]:
texts = [PREFIX + f"<Instruct>: {INSTRUCTION}\n<Query>: {query}\n<Document>: {p}" + SUFFIX
for p in products]
enc = tok(texts, return_tensors="pt", padding=True, truncation=True,
max_length=4096, add_special_tokens=False).to(model.device)
logits = model(**enc).logits[:, -1, :]
return torch.softmax(logits[:, [NO, YES]].float(), dim=-1)[:, 1].tolist()
print(score("jumpsuit under $100 boho", [
"title: Boho Wide-Leg Jumpsuit | price: $44 | color: rust",
"title: Bohemian Silk Jumpsuit | price: $239 | color: ivory",
]))
Product text may be the structured attribute schema or natural-language prose. The model was trained on the structured schema and carries its gain to prose (see the table).
Repository contents. config.json, model*.safetensors and the tokenizer files are
the unmodified Qwen3-Reranker-4B release; adapter_config.json and
adapter_model.safetensors are the ZooWork-ShopRanker LoRA adapter. The adapter is kept
separate rather than merged so the published model reproduces the paper numbers exactly.
If your serving stack needs a single set of weights, merging the adapter into the base in
bf16 stays within 0.4 points of the reported accuracy.
Training
LoRA (rank 16, α=32, dropout 0.05, on q/k/v/o projections) over judge-labelled preference pairs from private search traffic, labelled by a three-family LLM judge panel in both presentation orders, mixed with general retrieval replay data. This model was distilled from the aligned 8B, then sharpened on judged pairs. The benchmark comes from a disjoint slice of traffic (zero query and zero exact-pair overlap).
Limitations
- Labels come from an LLM judge panel and are not human-verified. No click, purchase or interleaving signal validates that the panel tracks what shoppers actually choose.
- Attribute hierarchy is not installed. On the AHP diagnostic, alignment lifts the
aggregate by transforming
priceandcolorwhile degradingstyleandproduct type; treat the gain as a side effect, not hierarchy-following. - Explicit price constraints are largely ignored by open rerankers, this one included.
- Zero-shot reasoning LLMs remain 8+ points more accurate overall, at far higher latency.
- The category mix leans toward apparel; evaluated in English only.
License and attribution
Apache License 2.0 (see LICENSE), inherited from Qwen/Qwen3-Reranker-4B
(Apache-2.0, by the Qwen team). This repository redistributes the base weights
and tokenizer unmodified; the ZooWork-ShopRanker modification is the LoRA adapter,
trained on judged preference data.
Citation
@article{xue2026zooworkshopranker,
title = {ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker},
author = {Xue, Siqiao and Liu, Shuxuan and Hu, Ning},
journal = {arXiv preprint arXiv:2609.31002},
year = {2026},
url = {https://arxiv.org/abs/2609.31002}
}
- Downloads last month
- 51