ZooWork-ShopRanker-4B

Project Page arXiv Model Models Dataset Code API ZooWork License

An e-commerce reranker aligned to judged shopper preference rather than topical relevance, from the paper ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker. This repository is self-contained: it holds the unmodified Qwen/Qwen3-Reranker-4B weights together with our LoRA adapter, so everything loads from this one repo id -- no separate base download. Same scoring interface as the base: one forward pass per (query, product), works at any k, and scores are cacheable per pair.

🚀 An optimized commercial version of the ZooWork-ShopRanker rerankers is available through the zoodata.ai API.

Results on ShopRank-Bench

Pairwise accuracy in percent on ShopRank-Bench (10,511 judged preference pairs from 2,991 queries); intervals are 95% query-clustered bootstrap CIs. Tiers are the number of judge families (of three) that committed to the label.

Model Product text Overall [95% CI] Gold (3/3) Silver (2/3) Bronze (1/3)
ZooWork-ShopRanker-4B Structured 81.2 [80.3, 82.0] 96.3 84.3 71.4
Qwen3-Reranker-4B (base) Structured 77.5 [76.5, 78.5] 92.7 81.4 66.7
ZooWork-ShopRanker-4B Natural language 79.5 [78.6, 80.4] 95.4 82.1 69.8
Qwen3-Reranker-4B (base) Natural language 77.9 [76.9, 78.9] 92.1 81.0 68.5

Significantly beats the strongest open baseline (Jina-reranker-m0: 79.2 structured, 78.1 natural language) in both formats and its own base by 3.7 points on structured text.

Report per tier: every reranker loses 20–25 points from gold to bronze, so an aggregate score is dominated by pairs nothing gets wrong.

Serving cost

One NVIDIA H200, fp16, 1024-token limit, 256 real (query, product) pairs, adapter merged into the weights (LoRA adds no inference cost once merged).

Params Batch-1 p50 Batch-1 p95 Docs/s (batch 32) Peak memory
4.02B 25.2 ms 33.3 ms 154.5 11.2 GB

Usage

Requires pip install transformers peft. Load the base weights and the adapter from this repo as shown: wrapping with PeftModel keeps the adapter in fp32, which reproduces the paper numbers exactly (loading the repo id with from_pretrained alone also attaches the adapter, but in bf16, and drifts by up to ~0.3 points).

The score is P(yes) = softmax over the bare no / yes logits at the last position of the native Qwen3-Reranker prompt -- exactly the scorer behind the numbers above.

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "srpone/zoowork-shopranker-4b"  # base weights + LoRA adapter, both in this repo
tok = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)
# Re-wrapping keeps the adapter in fp32, as in the paper. peft may warn "Already found a
# `peft_config` attribute"; that is expected here -- exactly one adapter ends up loaded.
model = PeftModel.from_pretrained(model, model_id).cuda().eval()

PREFIX = ("<|im_start|>system\nJudge whether the Document meets the requirements based on "
          "the Query and the Instruct provided. Note that the answer can only be \"yes\" "
          "or \"no\".<|im_end|>\n<|im_start|>user\n")
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCTION = "Given a web search query, retrieve relevant passages that answer the query"
YES, NO = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")

@torch.no_grad()
def score(query: str, products: list[str]) -> list[float]:
    texts = [PREFIX + f"<Instruct>: {INSTRUCTION}\n<Query>: {query}\n<Document>: {p}" + SUFFIX
             for p in products]
    enc = tok(texts, return_tensors="pt", padding=True, truncation=True,
              max_length=4096, add_special_tokens=False).to(model.device)
    logits = model(**enc).logits[:, -1, :]
    return torch.softmax(logits[:, [NO, YES]].float(), dim=-1)[:, 1].tolist()

print(score("jumpsuit under $100 boho", [
    "title: Boho Wide-Leg Jumpsuit | price: $44 | color: rust",
    "title: Bohemian Silk Jumpsuit | price: $239 | color: ivory",
]))

Product text may be the structured attribute schema or natural-language prose. The model was trained on the structured schema and carries its gain to prose (see the table).

Repository contents. config.json, model*.safetensors and the tokenizer files are the unmodified Qwen3-Reranker-4B release; adapter_config.json and adapter_model.safetensors are the ZooWork-ShopRanker LoRA adapter. The adapter is kept separate rather than merged so the published model reproduces the paper numbers exactly. If your serving stack needs a single set of weights, merging the adapter into the base in bf16 stays within 0.4 points of the reported accuracy.

Training

LoRA (rank 16, α=32, dropout 0.05, on q/k/v/o projections) over judge-labelled preference pairs from private search traffic, labelled by a three-family LLM judge panel in both presentation orders, mixed with general retrieval replay data. This model was distilled from the aligned 8B, then sharpened on judged pairs. The benchmark comes from a disjoint slice of traffic (zero query and zero exact-pair overlap).

Limitations

  • Labels come from an LLM judge panel and are not human-verified. No click, purchase or interleaving signal validates that the panel tracks what shoppers actually choose.
  • Attribute hierarchy is not installed. On the AHP diagnostic, alignment lifts the aggregate by transforming price and color while degrading style and product type; treat the gain as a side effect, not hierarchy-following.
  • Explicit price constraints are largely ignored by open rerankers, this one included.
  • Zero-shot reasoning LLMs remain 8+ points more accurate overall, at far higher latency.
  • The category mix leans toward apparel; evaluated in English only.

License and attribution

Apache License 2.0 (see LICENSE), inherited from Qwen/Qwen3-Reranker-4B (Apache-2.0, by the Qwen team). This repository redistributes the base weights and tokenizer unmodified; the ZooWork-ShopRanker modification is the LoRA adapter, trained on judged preference data.

Citation

@article{xue2026zooworkshopranker,
  title   = {ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker},
  author  = {Xue, Siqiao and Liu, Shuxuan and Hu, Ning},
  journal = {arXiv preprint arXiv:2609.31002},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.31002}
}
Downloads last month
51
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for srpone/zoowork-shopranker-4b

Adapter
(3)
this model

Dataset used to train srpone/zoowork-shopranker-4b

Collection including srpone/zoowork-shopranker-4b

Paper for srpone/zoowork-shopranker-4b