Introducing BreadBowl-Embed · Open weights · Apache-2.0
One vector is too few. One per token is too many.
BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text.
01 — The problem
High-precision retrieval reads your documents twice.
Once to find them, and again to decide which ones actually matter.
Modern search and RAG pipelines run in two stages. First, a bi-encoder compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance (the date, the exception, who did what to whom), a cross-encoder re-reads the top candidates next to the query and scores them again.
That second stage is where the precision comes from, and it is expensive in a very specific way: the reranker runs a full transformer pass over every (query, candidate) pair, every time a query arrives. Most of that work can't be cached ahead of time, because it depends on the query.
Agents make this worse. A coding or research agent doesn't issue one query per task; it issues dozens, and each one pays the reranking bill again. In my own coding-agent sessions, I'd estimate roughly 80% of the work is finding the right context: searching, reading and deciding what matters. Finding context is becoming the inner loop of AI work.
So we asked a simple question: how much of the reranker's job can the retrieval representation do by itself, using only what was computed and stored before the query arrived?
02 — The representation spectrum
One vector is too few. One per token is too many.
There have been two classic answers to the question how should a document be stored for search?
One vector per document. Bi-encoders such as Qwen3-Embedding, E5 or EmbeddingGemma pool the whole text into a single point. It's cheap to store and fast to search. But every query is compared against the same summary: a question about the date and a question about the architect hit the same point in space. The score is one dot product, so there is nothing left to refine. That's why single vectors so often get paired with a reranker.
One vector per token. Late-interaction models like ColBERT keep a contextual vector for every token and let each query token find its best match (MaxSim). That preserves detail and ranks well. The costs come with it: the index grows with every token stored, scoring cost grows with both query and document length, and all the scorer can do with a stored token vector is measure its similarity to a query token. Each query token takes its single best match; there is no separate content to read.
BreadBowl-Embed sits in between on purpose. Every passage gets a fixed budget of 16 slots, however long it is. Each slot carries two vectors with two different jobs:
- a routing vector (256-d) that says where to look. It is what gets indexed and searched.
- a value vector (256-d) that holds what to read once a query has decided where to look.
What do you actually need from a representation?
Here's Table 1 from our paper, made interactive. Switch on the requirements one at a time and watch which designs survive.
| Architecture | Retriever | Reranker | Extra backbone pass | Separate value readout | Stored doc. vectors | Scoring cost |
|---|---|---|---|---|---|---|
| Bi-encoder | ✓ | ✓* | None | — | 1 | O(h) |
| Cross-encoder | — | ✓ | Per pair | — | — | B(nq + nd) |
| Poly-encoder | — | ✓ | None | — | 1 | O(mq h) |
| ColBERT | ✓ | ✓ | None | — | nd | O(nq nd h) |
| ConstBERT | ✓ | ✓ | None | — | md | O(nq md h) |
| MVA (pooled) | — | ✓ | None | ✓ | 2md | O(mq md h) |
| BreadBowl-Embed | ✓ | ✓ | None | ✓ | 2md | O(mq md h) |
All 7 designs shown. Pick a requirement.
Switch on all four and exactly one row is left; every other design misses at least one. Bi-encoders and ColBERT can retrieve, but their stored vectors can only be matched, never read. ConstBERT fixes ColBERT's storage growth but keeps MaxSim. Multi-Vector Attention (MVA) introduced the key–value readout we build on, but its evaluated pipeline only reranks candidates that another system retrieved. Cross-encoders remain the gold standard for precision, but they can't retrieve and must re-read text.
BreadBowl-Embed's central idea is small to state: make the vectors that find a document the same vectors that decide how to read it. We call it polymorphic routing.
03 — How it works · Encode once
Sixteen learned questions, asked of every passage.
BreadBowl-Embed's Slot Encoder starts with an off-the-shelf language model: Qwen3.5-0.8B-Base reads the text and produces a contextual vector for every token. A ColBERT-style model would store those token vectors. We don't.
Instead, it pools them with Referenced Cross Attention (RCA): a bank of 16 learned reference vectors, shared by every input, each attends over all the token states and pulls out one slot. You can think of the references as 16 learned questions the model asks of every passage. The questions are fixed; the answers depend on the text.
Two small heads then read each slot: a linear routing head, and a value head (linear, plus a small residual MLP for extra capacity). Both run once, at indexing time, and their outputs are stored side by side: 16 × 256 routing numbers and 16 × 256 value numbers per passage. That's 8,192 numbers whether the passage is 20 tokens long or 384.
Retrieve relevant documents.), encoded by the same weights. There is no reference skip connection, feed-forward block or iterative slot refinement; the extractor is one pooling step.What do the slots actually hold?
Below is the query–document pair we inspect in the paper's appendix. Pick any slot to see which tokens it pooled. Query slot 13 puts nearly half its weight on completed. Document slot 13 pools the date: the last digit of 1889, then Fair, then the 7 of 1887. Slot 14 grabs Rome. Nothing in training asks for slots like these; it's simply what this example looks like.
Query slotsslot 13
Document slotsslot 13
04 — Retrieve with routing
Search with sixteen keys.
At query time, the query goes through the same encoder: one backbone pass over about twenty tokens. Each of its 16 routing vectors then searches an index holding every stored document routing vector. Hits are grouped by document into a shortlist, and the shortlist is rescored with the full routing score:
In words: for each query slot, take a soft maximum over the document's 16 slots, then average over query slots. As τr → 0 this becomes mean MaxSim, the late-interaction score familiar from ColBERT, just over 16 slots instead of hundreds of tokens. In the paper's pipeline each query slot pulls its top 128 slot hits, those are grouped into a 1,024-document shortlist, and the best 100 by routing score go on to reranking.
05 — Rerank by reading values
The same vectors get a second job.
Here's the step a single-vector model can't take. For each candidate we already have its 16 × 16 routing similarities. Soften them with a higher temperature and they become attention weights over the candidate's stored value vectors:
Each query slot gets its own readout: a weighted blend of the document's values, chosen by that query slot. The readout is compared with the query slot's own value vector, and the agreement is added to the routing score. There are no new projections at comparison time, no backbone and no text, just 16 × 16 attention over vectors that were already sitting in storage. The cost is independent of how long the candidate is.
The document never changes. What the query reads from it does.
What this question readsall query slots
Ask about the Colosseum, and all 16 query slots put their heaviest weight on document slot 9, the one that pooled Colosseum … is in Rome. Ask which museum is in Paris, and all 16 move to slot 5, which pooled the Louvre sentence. The two Eiffel Tower questions mostly read the same tower slots (11 and 12), because they're about the same thing; the difference shows up in individual query slots. In the completed question, two of the query slots that pooled completed lean hardest on slot 13, the date slot. It's the same 8,192 stored numbers every time.
A single-vector model can't do this: whatever you ask, it compares against the same point. ColBERT matches different tokens for different questions, but a match is all it can measure. BreadBowl-Embed adds a second channel: routing decides where to look, and values carry what's found there.
06 — Does reading the values help?
Same candidates. Better order.
The cleanest test of the idea is to take one trained model and one candidate list, then rank the list two ways: routing alone (QK), or routing plus the value readout (QKV). Recall is identical by construction, so any change in ranking comes from reading the values.
Show as a table
The gains are not uniform. The largest come on question answering and fact-checking: NQ (+11.9), TREC-COVID (+11.9, on only 50 queries), FEVER (+8.7) and Climate-FEVER (+6.9). The largest losses come where relevance means a similar or opposing text rather than contains the answer: ArguAna's counter-argument retrieval (−5.2) and Quora's duplicate questions (−2.2). That fits the mechanism, but the pattern isn't clean: SciFact, also claim verification, slips (−1.05), and Touché-2020, argument retrieval like ArguAna, gains (+5.8). There's also a confound. NQ, FEVER, HotpotQA and MS MARCO have training splits in our distillation data, so BEIR isn't uniformly zero-shot for this model; TREC-COVID's +11.9 is the largest gain on a task with no training split there.
Where the top result changes
A macro average can hide churn, so here's the blunt version. On our development set (5,000 queries searched over 3.41M documents), we counted the queries where reading the values moved a labeled passage into first place, and the ones where it pushed one out. This tally comes from the same runs as the paper's Tables 3 and 5, but it isn't in the paper.
Show as a table
Overall, values put a labeled passage on top 479 times and knocked one off 238 times: a net gain of 241, with plenty of churn. FEVER is lopsided (178 fixes, 7 breaks). MS MARCO goes the other way (112 fixes, 126 breaks), PubMedQA is slightly negative (2 vs 4), and SQuAD v2 breaks even. By nDCG@10, values improve six of the nine sources (paper, Table 5).
How close is that to a real reranker?
We trained the value pathway by distilling Qwen3-Reranker-0.6B, a cross-encoder. On the fixed development candidate lists used for checkpoint selection, the teacher scores 69.80 nDCG@10, routing alone 60.36, and routing plus values 63.40. So on those lists the readout closes about a third of the gap to its teacher, without running a transformer pass at rerank time.
Source-macro nDCG@10 × 100 on the saved development candidates used to select the checkpoint (step 4,000), against one 0.6B teacher. The readout recovers 3.04 of the 9.44-point gap, about 32%.
It doesn't replace a cross-encoder: two-thirds of that gap remains, and you can still run one on top. What it does is move a real share of the reranker's value into a step that costs about 0.27 MFLOPs per candidate.
Try it on real queries
These are real questions and fact-checking claims from an earlier validation set, each with a 10–12 passage shortlist taken from saved candidate lists and re-scored with the released weights. Flip between routing-only and routing-plus-values ordering and watch where the labeled passage lands. The last one is a case where values make things worse, because that happens too.
Look at what routing alone puts near the top of the Natural Questions lists: title stubs and list fragments such as Thursday Night Football or 1. “Take It Easy”. They match the topic, but there's nothing in them that answers the question, and their value scores come out negative, so they sink. The Hamilton claim shows another pattern: routing ranks several other Alexander Hamiltons first, and the values put the Founding Father's biography, the labeled evidence, on top. The miss is instructive too. For “The Rocket”, the values latch onto an answer-shaped passage (a roller coaster called The Rocket that was destroyed in 1979) about the wrong Rocket.
Routing finds the topic. Values look for the answer.
07 — Training
Teach it to rank. Replay how to find.
Sharing one representation between retrieval and reranking sets up a tug-of-war. Training the value pathway to imitate a cross-encoder updates the shared Slot Encoder, and that can quietly move the routing vectors that retrieval depends on. A sharper reranker over worse candidates isn't progress.
Learn to find, and lay down the slots
Backbone, slot extractor and both heads train end to end on weak query–document pairs from LightOn's curated pre-training collection (a 30M-pair preparation target). One loss teaches routing to pick out positives among in-batch documents; a second teaches the combined score to separate positives from routing-mined hard negatives; a diversity penalty discourages the 16 routing slots from collapsing onto each other.
8 × H200 · about 8 hoursLearn to read, without forgetting how to find
Freeze the backbone and routing head. Distill the teacher's listwise ranking (KL divergence over 32 candidates per query, 500,000 queries from nine sources) into the slot extractor and value modules. Every eighth update, replay a contrastive routing loss on labeled query–positive pairs, so candidate discovery stays supervised while relevance refinement improves.
8 × RTX 5090 · about 10 hoursDoes replay matter? We ran the same distillation with and without it: same initialization, same queries, 4,000 updates each, one seed per arm. Replay does add training compute (8.2M extra query–positive presentations). Each checkpoint then re-encoded the development corpora and retrieved its own candidates, so the comparison captures discovery as well as ordering.
| Training | Recall@100 | nDCG@10 routing | nDCG@10 routing + values |
|---|---|---|---|
| KL distillation only | 86.63 | 58.64 | 62.57 |
| KL + retrieval replay | 87.19 +0.56 | 60.40 +1.76 | 63.42 +0.85 |
Source-macro averages over the 5,000-query development set (paper, Table 3). The replay gain in final nDCG@10 (0.84 before rounding) has a 95% bootstrap interval of [0.45, 1.24] points.
Replay improves both what we find and how we order it. One honest wrinkle: recall of the coarse shortlist, before routing picks the final 100, dipped slightly (88.50 → 88.38), so replay doesn't preserve every aspect of candidate coverage.
08 — Where it stands
Not the leaderboard leader. A new shape of model.
We want to be precise about what this release is and isn't. On published BEIR averages, BreadBowl-Embed trails strong single-vector models, including some much smaller ones:
Show as a table
Because published numbers aren't directly comparable, we also re-ran two of those models ourselves on identical development data, with the same queries, corpora, labels and token limits:
| Model | nDCG@10 | Recall@100 |
|---|---|---|
| Qwen3-Embedding-0.6B | 62.25 | 87.42 |
| Jina v5 text-small (retrieval) | 66.14 | 90.73 |
| BreadBowl-Embed, routing only | 60.40 | 87.19 |
| BreadBowl-Embed, routing + values | 63.42 | 87.19 |
Source-macro averages over nine development sources (paper, Table 9). Each system retrieves its own candidates. Training data, model size and training compute are not matched. These queries come from the same sources as BreadBowl-Embed's training data, and the set was used for model selection.
Routing alone trails Qwen3-Embedding-0.6B. Reading the values puts BreadBowl-Embed ahead on average, at similar recall, but the lead is concentrated: it wins four of the nine sources, and FEVER and NQ carry the average. Jina v5 text-small is ahead of both.
So why release now? Because the interesting result isn't the leaderboard position; it's the mechanism. Same model, same candidates, +3.07 nDCG@10 on BEIR, from information that was already sitting in the index. (That comparison switches the value score off at inference in one trained model; it isn't a separately trained routing-only model.) A single-vector index has no stored content to read at query time, so it has no equivalent step to take. Whether the value gain grows or shrinks as routing improves is an open question: in our replay ablation, better routing came with a smaller gain from values (3.93 → 3.02 points). And we haven't scaled any part of the recipe yet.
- End-to-end latency and throughput against a retriever + cross-encoder stack. The efficiency argument here is architectural arithmetic, not a benchmark.
- Cross-encoder-level precision. On our development lists, the 0.6B teacher is still 6.4 nDCG@10 points ahead.
- A head-to-head with ColBERT-family models under matched training.
- Variance across training seeds, or a full audit of overlap between training data and BEIR. Some BEIR domains appear in our adaptation data, and earlier SciFact and NFCorpus runs informed development.
- Compatibility of stored embeddings across model versions.
09 — Try it
Encode once. Retrieve and rerank.
The weights are on Hugging Face under Apache-2.0 as breadbowl-embed-v1.2-preview,6 and the inference toolkit is on GitHub (it isn't on PyPI yet). It runs on CPU by default, and on Apple Silicon with device="mps". Python 3.12 or newer.
python -m pip install "git+https://github.com/BreadBowlAI/breadbowl-embed.git"
from breadbowl_embed import BreadBowl
model = BreadBowl.from_pretrained("dotproductx/breadbowl-embed-v1.2-preview")
passages = [
"Octopuses have three hearts.",
"The Pacific is Earth's largest ocean.",
"Plants use sunlight during photosynthesis.",
]
for hit in model.rerank("How many hearts does an octopus have?", passages):
print(hit.score, passages[hit.document_index])
For repeated queries, encode documents once and reuse the stored slots. The index returns routing, value and combined scores for every hit:
documents = model.encode_documents(passages, batch_size=2) # routing + value slots, stored once
index = model.index(documents, ids=["octopus", "ocean", "plants"])
queries = model.encode_queries(["How many hearts does an octopus have?"])
for hit in index.search(queries, top_k=2, candidate_k=3)[0]:
print(hit.id, hit.routing_score, hit.value_score, hit.score)
print(documents.routing.shape) # [3, 16, 256]
print(documents.value.shape) # [3, 16, 256]
The bundled index is an exact linear scan over every document's routing slots, meant for experiments. It doesn't implement the paper's per-slot shortlist pipeline, and it has no ANN backend. Raw vectors take 32 KiB per passage in float32. Keep document values unnormalized before the attention read; normalizing them changes the scorer.
10 — Why we're building this
Representations that can be read, not just matched.
More and more AI products are retrieval products underneath. Coding agents search repositories, enterprise copilots search company knowledge, research agents search the web, and recommendation systems search through people, products and intent. Underneath all of them sits a representation layer that, today, mostly points at nearest neighbors and leaves the real judgment to a second model.
BreadBowl's bet is that this layer should carry more structure: enough to rank, to adapt to a domain, and to be read differently by every question. BreadBowl-Embed is the first step: one stored representation that both finds and reads.
Want to see what it finds in your data?
Alpha access to BreadBowl's hosted service is open for model evaluations, design partnerships and teams building retrieval-heavy products. We'd especially love to hear from you if a reranker is your bottleneck.
contact@breadbowl.aiCitation
Notes
- Cross-encoder estimate: forward FLOPs ≈ 2 × non-embedding parameters × tokens. For Qwen3-Reranker-0.6B (our distillation teacher, ≈0.44B non-embedding parameters) on a ≈420-token (query + 384-token passage) pair, that is ≈0.4 TFLOPs per candidate once attention is included, or ≈40 TFLOPs for 100 candidates; the reranker's prompt template adds tokens, so this is conservative. Shorter passages shrink it proportionally. BreadBowl-Embed's value read: 2 × (16·16·256 + 16·16·256 + 16·256) ≈ 0.27 MFLOPs per candidate, or ≈27 MFLOPs for 100. Its first stage is heavier than a single-vector search, though: an exact flat search compares 16 query vectors with 16 vectors per document (65,536 multiply-adds per document, 64× a 1,024-d single vector), and rescoring a 1,024-document shortlist with full routing scores adds ≈0.13 GFLOPs. None of this includes moving stored vectors (16–32 KiB per candidate). We have not measured end-to-end latency.
- Counts follow Table 1 of the paper and precede compression and index duplication. ColBERT is shown with 128-d token vectors and 32-token queries; BreadBowl-Embed with 16 slots of 256 routing + 256 value dimensions. Scoring cost: ColBERT 32 × n × 128 multiply-adds; BreadBowl-Embed 16 × 16 × 256 (routing) + 16 × 16 × 256 (value read) + 16 × 256 (comparison) per passage. At 384 tokens a 2-bit ColBERTv2 index needs ≈13.5 KiB per passage, versus 16 KiB for BreadBowl-Embed's slots in bf16.
- Demo numbers come from running the released weights through a NumPy re-implementation of the inference path, in float32 on CPU. Its tokenization matches the reference token IDs, and its routing vectors agree with the reference GPU outputs (bf16) at a per-slot cosine of 0.9998 or better, so values can differ from the paper's figures in the third decimal.
- Official BEIR full corpora and relevance judgments, test splits except MS MARCO (dev). nDCG computed with pytrec_eval. CQADupStack's 12 forums are averaged into one task before the 15-task macro average. Candidates come from Algorithm 1 of the paper (128 slot hits per query slot, a 1,024-document shortlist, the top 100 by routing score) with exact flat search. One fixed checkpoint and one set of retrieval settings for every task; results are point estimates.
- 5,000 development queries from nine sources (LightOn subsets of MS MARCO, NQ, TriviaQA, SQuAD v2, HotpotQA and FEVER, plus StackExchange duplicates, SciRepEval search and PubMedQA), each searched over its own prepared corpus (3.41M documents in total). They were held out before candidate mining but come from the same sources as the training data, and they were used for model selection. The top-1 fix/break counts are our own tally from the same runs.
- The checkpoint is published as
breadbowl-embed-v1.2-preview. "v1.2" is the internal research iteration; this is BreadBowl's first public model release.