Introducing BreadBowl-Embed  ·  Open weights  ·  Apache-2.0

One vector is too few. One per token is too many.

BreadBowl-Embed is a new late-interaction architecture for embeddings. Instead of one vector per document, or one per token, it stores every passage as 16 routing–value slots. The routing vectors find candidates in an index; then the same vectors decide which stored values each query reads. Retrieval and reranking share one representation: documents are encoded once, and reranking never re-reads their text.

Ming (Jerry) Xu Founder, BreadBowl AI October 2026 15 min read
One passage, encoded onceReal model weights
16 stored slotsrouting ▴ / value ▾
Routing vector · where to lookValue vector · what to read Hover or tap a slot to see what it pooled.
The Slot Encoder reads this passage once and pools it into 16 slots. The flowing lines show the strongest of the released checkpoint's actual pooling weights for this text (mean of four heads). Each slot is stored as a routing vector and a value vector; that's everything the index keeps.

01 — The problem

High-precision retrieval reads your documents twice.

Once to find them, and again to decide which ones actually matter.

Modern search and RAG pipelines run in two stages. First, a bi-encoder compresses every document into a single vector, so a nearest-neighbor index can search millions of them in milliseconds. Then, because one vector throws away the details that decide relevance (the date, the exception, who did what to whom), a cross-encoder re-reads the top candidates next to the query and scores them again.

That second stage is where the precision comes from, and it is expensive in a very specific way: the reranker runs a full transformer pass over every (query, candidate) pair, every time a query arrives. Most of that work can't be cached ahead of time, because it depends on the query.

Agents make this worse. A coding or research agent doesn't issue one query per task; it issues dozens, and each one pays the reranking bill again. In my own coding-agent sessions, I'd estimate roughly 80% of the work is finding the right context: searching, reading and deciding what matters. Finding context is becoming the inner loop of AI work.

So we asked a simple question: how much of the reranker's job can the retrieval representation do by itself, using only what was computed and stored before the query arrived?

One query, two retrieval stacksBack-of-envelope arithmetic
Embed + cross-encoder rerankThe usual two-stage stack
Query~20 tokens
Embed query1 transformer pass
Vector indexnearest neighbors
100 candidatestext fetched again
Cross-encoder1 pass per candidate
Ranked top 10
0transformer passes 0candidate tokens re-read 0rerank FLOPs
BreadBowl-EmbedRetrieve and rerank from stored slots
Query~20 tokens
Slot Encoder1 transformer pass
Routing index16 keys per query
100 candidatesstored slots loaded
Read stored values16×16 attention, no text
Ranked top 10
0transformer passes 0candidate tokens re-read 0rerank FLOPs
One transformer pass over (query + candidate) textAttention over stored vectors
Each orange square is one cross-encoder pass over a (query, candidate) pair, assuming 384-token passages. BreadBowl-Embed's rerank step is 16×16 attention over vectors written at indexing time. Both stacks still encode the query once and search an index, and BreadBowl-Embed's 16-vector search does more arithmetic than a single-vector one. Compute figures are arithmetic estimates, not latency measurements; the assumptions are in note 1.

02 — The representation spectrum

One vector is too few. One per token is too many.

There have been two classic answers to the question how should a document be stored for search?

One vector per document. Bi-encoders such as Qwen3-Embedding, E5 or EmbeddingGemma pool the whole text into a single point. It's cheap to store and fast to search. But every query is compared against the same summary: a question about the date and a question about the architect hit the same point in space. The score is one dot product, so there is nothing left to refine. That's why single vectors so often get paired with a reranker.

One vector per token. Late-interaction models like ColBERT keep a contextual vector for every token and let each query token find its best match (MaxSim). That preserves detail and ranks well. The costs come with it: the index grows with every token stored, scoring cost grows with both query and document length, and all the scorer can do with a stored token vector is measure its similarity to a query token. Each query token takes its single best match; there is no separate content to read.

BreadBowl-Embed sits in between on purpose. Every passage gets a fixed budget of 16 slots, however long it is. Each slot carries two vectors with two different jobs:

  • a routing vector (256-d) that says where to look. It is what gets indexed and searched.
  • a value vector (256-d) that holds what to read once a query has decided where to look.
What gets stored for one passageArithmetic from the architectures
384 tokens
Single vectorBi-encoder, e.g. 1,024-d
One vector per tokenColBERT-style, 128-d per token
BreadBowl-Embed16 slots × (256 routing + 256 value)
Numbers stored are counted before compression. In bytes, a compressed ColBERTv2 index (roughly 20–36 bytes per token vector) is currently smaller than our uncompressed slots (32 KiB per passage in float32, 16 KiB in bf16); the structural difference is that ColBERT's vector count and scoring cost grow with length. Scoring cost is multiply-adds per candidate, with 32-token ColBERT queries and 16 query slots for BreadBowl-Embed. BreadBowl-Embed reads at most 384 tokens and truncates the rest, so the widget assumes long documents are split into passages first. Details in note 2.

What do you actually need from a representation?

Here's Table 1 from our paper, made interactive. Switch on the requirements one at a time and watch which designs survive.

Require:
Architecture Retriever Reranker Extra backbone pass Separate value readout Stored doc. vectors Scoring cost
Bi-encoder✓✓*None—1O(h)
Cross-encoder—✓Per pair——B(nq + nd)
Poly-encoder—✓None—1O(mq h)
ColBERT✓✓None—ndO(nq nd h)
ConstBERT✓✓None—mdO(nq md h)
MVA (pooled)—✓None✓2mdO(mq md h)
BreadBowl-Embed✓✓None✓2mdO(mq md h)

All 7 designs shown. Pick a requirement.

nq, nd: query and document token counts. mq, md: pooled vectors per input (16 for BreadBowl-Embed). h: vector width. B(ℓ): one backbone pass over ℓ tokens. *A bi-encoder "reranks" by reusing the same dot product it retrieved with. Costs describe candidate scoring only. From Table 1 of the paper.

Switch on all four and exactly one row is left; every other design misses at least one. Bi-encoders and ColBERT can retrieve, but their stored vectors can only be matched, never read. ConstBERT fixes ColBERT's storage growth but keeps MaxSim. Multi-Vector Attention (MVA) introduced the key–value readout we build on, but its evaluated pipeline only reranks candidates that another system retrieved. Cross-encoders remain the gold standard for precision, but they can't retrieve and must re-read text.

BreadBowl-Embed's central idea is small to state: make the vectors that find a document the same vectors that decide how to read it. We call it polymorphic routing.

03 — How it works · Encode once

Sixteen learned questions, asked of every passage.

BreadBowl-Embed's Slot Encoder starts with an off-the-shelf language model: Qwen3.5-0.8B-Base reads the text and produces a contextual vector for every token. A ColBERT-style model would store those token vectors. We don't.

Instead, it pools them with Referenced Cross Attention (RCA): a bank of 16 learned reference vectors, shared by every input, each attends over all the token states and pulls out one slot. You can think of the references as 16 learned questions the model asks of every passage. The questions are fixed; the answers depend on the text.

Two small heads then read each slot: a linear routing head, and a value head (linear, plus a small residual MLP for extra capacity). Both run once, at indexing time, and their outputs are stored side by side: 16 × 256 routing numbers and 16 × 256 value numbers per passage. That's 8,192 numbers whether the passage is 20 tokens long or 384.

The Slot EncoderSame weights for queries and documents
Textquery (≤ 128 tokens) or passage≤ 384 tokens
BackboneQwen3.5-0.8B-Base, 24 layersn × 1024
RCA pooling16 learned references attend over all tokens (4 heads)16 × 1024
Routing headlinear16 × 256
Value headlinear + residual MLP16 × 256
Stored onceindex the routing rows; keep the values beside them8,192 numbers
A query is just a short input with an instruction prefix (Retrieve relevant documents.), encoded by the same weights. There is no reference skip connection, feed-forward block or iterative slot refinement; the extractor is one pooling step.

What do the slots actually hold?

Below is the query–document pair we inspect in the paper's appendix. Pick any slot to see which tokens it pooled. Query slot 13 puts nearly half its weight on completed. Document slot 13 pools the date: the last digit of 1889, then Fair, then the 7 of 1887. Slot 14 grabs Rome. Nothing in training asks for slots like these; it's simply what this example looks like.

Slot pooling, query and documentReal model outputs

Query slotsslot 13

Document slotsslot 13

Highlight strength is the slot's pooling weight on each token (mean of 4 heads); tokens with at least 8% of the weight show percentages. Selecting a query slot also outlines the document slot whose routing vector is closest to it (cosine). These weights pool contextual states, so they aren't word-importance scores: the last digit of a year can carry the whole year. One example doesn't establish stable slot roles. Computed with the released weights (note 3).

04 — Retrieve with routing

Search with sixteen keys.

At query time, the query goes through the same encoder: one backbone pass over about twenty tokens. Each of its 16 routing vectors then searches an index holding every stored document routing vector. Hits are grouped by document into a shortlist, and the shortlist is rescored with the full routing score:

Slot similarities C=R^qR^d⊤(16 × 16 cosines)
Routing score sr=1m∑i=1mτrlog(1m∑j=1meCij/τr),τr=0.02

In words: for each query slot, take a soft maximum over the document's 16 slots, then average over query slots. As τr → 0 this becomes mean MaxSim, the late-interaction score familiar from ColBERT, just over 16 slots instead of hundreds of tokens. In the paper's pipeline each query slot pulls its top 128 slot hits, those are grouped into a 1,024-document shortlist, and the best 100 by routing score go on to reranking.

05 — Rerank by reading values

The same vectors get a second job.

Here's the step a single-vector model can't take. For each candidate we already have its 16 × 16 routing similarities. Soften them with a higher temperature and they become attention weights over the candidate's stored value vectors:

Attention from routing A=softmaxj(C/τv),τv=0.1
Read stored values U=AVd(one readout per query slot)
Value score sv=1m∑i=1mcos(viq,ui)
Final score s=0.75sr+0.25sv

Each query slot gets its own readout: a weighted blend of the document's values, chosen by that query slot. The readout is compared with the query slot's own value vector, and the agreement is added to the routing score. There are no new projections at comparison time, no backbone and no text, just 16 × 16 attention over vectors that were already sitting in storage. The cost is independent of how long the candidate is.

The document never changes. What the query reads from it does.

Same stored document, four different questionsReal model outputs

What this question readsall query slots

Routing score–
Value score–
Final 0.75·r + 0.25·v–
Rows: query slots · click one to isolate itColumns: stored document slots

Heatmap: value-attention weights A for the selected question (stronger color means more weight). Highlighted text: the document tokens pooled into the slots this question reads (attention weights multiplied by pooling weights). The stored routing and value vectors are identical in every view; only the query changes. Computed with the released weights (note 3).

Ask about the Colosseum, and all 16 query slots put their heaviest weight on document slot 9, the one that pooled Colosseum … is in Rome. Ask which museum is in Paris, and all 16 move to slot 5, which pooled the Louvre sentence. The two Eiffel Tower questions mostly read the same tower slots (11 and 12), because they're about the same thing; the difference shows up in individual query slots. In the completed question, two of the query slots that pooled completed lean hardest on slot 13, the date slot. It's the same 8,192 stored numbers every time.

A single-vector model can't do this: whatever you ask, it compares against the same point. ColBERT matches different tokens for different questions, but a match is all it can measure. BreadBowl-Embed adds a second channel: routing decides where to look, and values carry what's found there.

06 — Does reading the values help?

Same candidates. Better order.

The cleanest test of the idea is to take one trained model and one candidate list, then rank the list two ways: routing alone (QK), or routing plus the value readout (QKV). Recall is identical by construction, so any change in ranking comes from reading the values.

48.47 → 51.54
BEIR macro nDCG@10 across 15 tasks, same 100 candidates per query
10 of 15
BEIR tasks improve when values are read
+11.9
nDCG@10 points on NQ, and again on TREC-COVID
≈ ⅓
of the gap to its 0.6B cross-encoder teacher closed, on the dev lists used for checkpoint selection
BEIR nDCG@10: routing only vs routing + valuesPaper, Table 2
Routing only (QK)Routing + values, higherRouting + values, lower
Show as a table
nDCG@10 × 100 on full BEIR corpora. QK and QKV rank the same 100 routing-selected candidates per query, so recall@100 is identical for both (macro 67.91). Δ is computed before rounding. Protocol in note 4.

The gains are not uniform. The largest come on question answering and fact-checking: NQ (+11.9), TREC-COVID (+11.9, on only 50 queries), FEVER (+8.7) and Climate-FEVER (+6.9). The largest losses come where relevance means a similar or opposing text rather than contains the answer: ArguAna's counter-argument retrieval (−5.2) and Quora's duplicate questions (−2.2). That fits the mechanism, but the pattern isn't clean: SciFact, also claim verification, slips (−1.05), and Touché-2020, argument retrieval like ArguAna, gains (+5.8). There's also a confound. NQ, FEVER, HotpotQA and MS MARCO have training splits in our distillation data, so BEIR isn't uniformly zero-shot for this model; TREC-COVID's +11.9 is the largest gain on a task with no training split there.

Where the top result changes

A macro average can hide churn, so here's the blunt version. On our development set (5,000 queries searched over 3.41M documents), we counted the queries where reading the values moved a labeled passage into first place, and the ones where it pushed one out. This tally comes from the same runs as the paper's Tables 3 and 5, but it isn't in the paper.

Top-1 result fixed vs broken by reading valuesFresh retrieval runs, dev set
Values knocked the right passage out of #1Values moved the right passage into #1
Show as a table
Counted over the 4,572 development queries whose labeled passage appears among the 100 routing candidates. Each source is retrieved from its own prepared corpus. The development set was also used for checkpoint selection, so read this as a diagnostic rather than a held-out test (note 5).

Overall, values put a labeled passage on top 479 times and knocked one off 238 times: a net gain of 241, with plenty of churn. FEVER is lopsided (178 fixes, 7 breaks). MS MARCO goes the other way (112 fixes, 126 breaks), PubMedQA is slightly negative (2 vs 4), and SQuAD v2 breaks even. By nDCG@10, values improve six of the nine sources (paper, Table 5).

How close is that to a real reranker?

We trained the value pathway by distilling Qwen3-Reranker-0.6B, a cross-encoder. On the fixed development candidate lists used for checkpoint selection, the teacher scores 69.80 nDCG@10, routing alone 60.36, and routing plus values 63.40. So on those lists the readout closes about a third of the gap to its teacher, without running a transformer pass at rerank time.

Gap to the cross-encoder teacherPaper, Table 7
60.36Routing only
63.40+ values
69.80Cross-encoder teacher

Source-macro nDCG@10 × 100 on the saved development candidates used to select the checkpoint (step 4,000), against one 0.6B teacher. The readout recovers 3.04 of the 9.44-point gap, about 32%.

It doesn't replace a cross-encoder: two-thirds of that gap remains, and you can still run one on top. What it does is move a real share of the reranker's value into a step that costs about 0.27 MFLOPs per candidate.

Try it on real queries

These are real questions and fact-checking claims from an earlier validation set, each with a 10–12 passage shortlist taken from saved candidate lists and re-scored with the released weights. Flip between routing-only and routing-plus-values ordering and watch where the labeled passage lands. The last one is a case where values make things worse, because that happens too.

Rerank playgroundReal queries · released weights

    Queries and passages come from the Natural Questions, FEVER and HotpotQA subsets of an earlier validation set (not the 5,000-query dev set). Each shortlist is the top of a saved candidate list built with an earlier checkpoint, re-scored here with the released weights (note 3); we haven't verified that these queries are disjoint from the training pool. The green tag marks the dataset's labeled passage. Bars show the released checkpoint's routing score sr and value score sv; the final score is 0.75·sr + 0.25·sv. These examples were picked to show the mechanism, so trust the aggregate numbers above over any single example.

    Look at what routing alone puts near the top of the Natural Questions lists: title stubs and list fragments such as Thursday Night Football or 1. “Take It Easy”. They match the topic, but there's nothing in them that answers the question, and their value scores come out negative, so they sink. The Hamilton claim shows another pattern: routing ranks several other Alexander Hamiltons first, and the values put the Founding Father's biography, the labeled evidence, on top. The miss is instructive too. For “The Rocket”, the values latch onto an answer-shaped passage (a roller coaster called The Rocket that was destroyed in 1979) about the wrong Rocket.

    Routing finds the topic. Values look for the answer.

    07 — Training

    Teach it to rank. Replay how to find.

    Sharing one representation between retrieval and reranking sets up a tug-of-war. Training the value pathway to imitate a cross-encoder updates the shared Slot Encoder, and that can quietly move the routing vectors that retrieval depends on. A sharper reranker over worse candidates isn't progress.

    Stage 1 · Weak contrastive training

    Learn to find, and lay down the slots

    Backbone, slot extractor and both heads train end to end on weak query–document pairs from LightOn's curated pre-training collection (a 30M-pair preparation target). One loss teaches routing to pick out positives among in-batch documents; a second teaches the combined score to separate positives from routing-mined hard negatives; a diversity penalty discourages the 16 routing slots from collapsing onto each other.

    8 × H200 · about 8 hours
    Stage 2 · Distillation with retrieval replay

    Learn to read, without forgetting how to find

    Freeze the backbone and routing head. Distill the teacher's listwise ranking (KL divergence over 32 candidates per query, 500,000 queries from nine sources) into the slot extractor and value modules. Every eighth update, replay a contrastive routing loss on labeled query–positive pairs, so candidate discovery stays supervised while relevance refinement improves.

    8 × RTX 5090 · about 10 hours
    Distillation update (values)+ retrieval replay (routing), every 8th update
    The two-stage recipe from the paper (§3.4 and Appendix A.3). The reported runs used about 64 H200 GPU-hours and 80 RTX 5090 GPU-hours, not counting development, data preparation, teacher scoring, or the earlier distillation run that stage 2 continues from (1,600 updates on 20,000 queries).

    Does replay matter? We ran the same distillation with and without it: same initialization, same queries, 4,000 updates each, one seed per arm. Replay does add training compute (8.2M extra query–positive presentations). Each checkpoint then re-encoded the development corpora and retrieved its own candidates, so the comparison captures discovery as well as ordering.

    TrainingRecall@100nDCG@10 routingnDCG@10 routing + values
    KL distillation only86.6358.6462.57
    KL + retrieval replay87.19 +0.5660.40 +1.7663.42 +0.85

    Source-macro averages over the 5,000-query development set (paper, Table 3). The replay gain in final nDCG@10 (0.84 before rounding) has a 95% bootstrap interval of [0.45, 1.24] points.

    Replay improves both what we find and how we order it. One honest wrinkle: recall of the coarse shortlist, before routing picks the final 100, dipped slightly (88.50 → 88.38), so replay doesn't preserve every aspect of candidate coverage.

    08 — Where it stands

    Not the leaderboard leader. A new shape of model.

    We want to be precise about what this release is and isn't. On published BEIR averages, BreadBowl-Embed trails strong single-vector models, including some much smaller ones:

    Published BEIR means (nDCG@10 × 100)Protocols differ · paper, Table 8
    Show as a table
    Published means and parameter counts from Akram et al. (2026), Table 4. Those reports don't control input formatting, truncation, checkpoint revisions or retrieval settings, so this is literature context rather than a matched comparison. The two BreadBowl-Embed rows are our own evaluation (paper, Table 2), and their parameter count is the nominal backbone size.

    Because published numbers aren't directly comparable, we also re-ran two of those models ourselves on identical development data, with the same queries, corpora, labels and token limits:

    ModelnDCG@10Recall@100
    Qwen3-Embedding-0.6B62.2587.42
    Jina v5 text-small (retrieval)66.1490.73
    BreadBowl-Embed, routing only60.4087.19
    BreadBowl-Embed, routing + values63.4287.19

    Source-macro averages over nine development sources (paper, Table 9). Each system retrieves its own candidates. Training data, model size and training compute are not matched. These queries come from the same sources as BreadBowl-Embed's training data, and the set was used for model selection.

    Routing alone trails Qwen3-Embedding-0.6B. Reading the values puts BreadBowl-Embed ahead on average, at similar recall, but the lead is concentrated: it wins four of the nine sources, and FEVER and NQ carry the average. Jina v5 text-small is ahead of both.

    So why release now? Because the interesting result isn't the leaderboard position; it's the mechanism. Same model, same candidates, +3.07 nDCG@10 on BEIR, from information that was already sitting in the index. (That comparison switches the value score off at inference in one trained model; it isn't a separately trained routing-only model.) A single-vector index has no stored content to read at query time, so it has no equivalent step to take. Whether the value gain grows or shrinks as routing improves is an open question: in our replay ablation, better routing came with a smaller gain from values (3.93 → 3.02 points). And we haven't scaled any part of the recipe yet.

    What we haven't shown yet
    • End-to-end latency and throughput against a retriever + cross-encoder stack. The efficiency argument here is architectural arithmetic, not a benchmark.
    • Cross-encoder-level precision. On our development lists, the 0.6B teacher is still 6.4 nDCG@10 points ahead.
    • A head-to-head with ColBERT-family models under matched training.
    • Variance across training seeds, or a full audit of overlap between training data and BEIR. Some BEIR domains appear in our adaptation data, and earlier SciFact and NFCorpus runs informed development.
    • Compatibility of stored embeddings across model versions.

    09 — Try it

    Encode once. Retrieve and rerank.

    The weights are on Hugging Face under Apache-2.0 as breadbowl-embed-v1.2-preview,6 and the inference toolkit is on GitHub (it isn't on PyPI yet). It runs on CPU by default, and on Apple Silicon with device="mps". Python 3.12 or newer.

    python -m pip install "git+https://github.com/BreadBowlAI/breadbowl-embed.git"
    from breadbowl_embed import BreadBowl
    
    model = BreadBowl.from_pretrained("dotproductx/breadbowl-embed-v1.2-preview")
    passages = [
        "Octopuses have three hearts.",
        "The Pacific is Earth's largest ocean.",
        "Plants use sunlight during photosynthesis.",
    ]
    for hit in model.rerank("How many hearts does an octopus have?", passages):
        print(hit.score, passages[hit.document_index])

    For repeated queries, encode documents once and reuse the stored slots. The index returns routing, value and combined scores for every hit:

    documents = model.encode_documents(passages, batch_size=2)   # routing + value slots, stored once
    index = model.index(documents, ids=["octopus", "ocean", "plants"])
    queries = model.encode_queries(["How many hearts does an octopus have?"])
    
    for hit in index.search(queries, top_k=2, candidate_k=3)[0]:
        print(hit.id, hit.routing_score, hit.value_score, hit.score)
    
    print(documents.routing.shape)  # [3, 16, 256]
    print(documents.value.shape)    # [3, 16, 256]

    The bundled index is an exact linear scan over every document's routing slots, meant for experiments. It doesn't implement the paper's per-slot shortlist pipeline, and it has no ANN backend. Raw vectors take 32 KiB per passage in float32. Keep document values unnormalized before the attention read; normalizing them changes the scorer.

    10 — Why we're building this

    Representations that can be read, not just matched.

    More and more AI products are retrieval products underneath. Coding agents search repositories, enterprise copilots search company knowledge, research agents search the web, and recommendation systems search through people, products and intent. Underneath all of them sits a representation layer that, today, mostly points at nearest neighbors and leaves the real judgment to a second model.

    BreadBowl's bet is that this layer should carry more structure: enough to rank, to adapt to a domain, and to be read differently by every question. BreadBowl-Embed is the first step: one stored representation that both finds and reads.

    Want to see what it finds in your data?

    Alpha access to BreadBowl's hosted service is open for model evaluations, design partnerships and teams building retrieval-heavy products. We'd especially love to hear from you if a reranker is your bottleneck.

    contact@breadbowl.ai

    Citation

    @misc{xu2026breadbowlembed, title = {BreadBowl-Embed: Encode Once, Retrieve and Rerank}, author = {Xu, Ming}, year = {2026}, note = {Preprint}, url = {https://huggingface.co/dotproductx/breadbowl-embed-v1.2-preview} }

    Notes

    1. Cross-encoder estimate: forward FLOPs ≈ 2 × non-embedding parameters × tokens. For Qwen3-Reranker-0.6B (our distillation teacher, ≈0.44B non-embedding parameters) on a ≈420-token (query + 384-token passage) pair, that is ≈0.4 TFLOPs per candidate once attention is included, or ≈40 TFLOPs for 100 candidates; the reranker's prompt template adds tokens, so this is conservative. Shorter passages shrink it proportionally. BreadBowl-Embed's value read: 2 × (16·16·256 + 16·16·256 + 16·256) ≈ 0.27 MFLOPs per candidate, or ≈27 MFLOPs for 100. Its first stage is heavier than a single-vector search, though: an exact flat search compares 16 query vectors with 16 vectors per document (65,536 multiply-adds per document, 64× a 1,024-d single vector), and rescoring a 1,024-document shortlist with full routing scores adds ≈0.13 GFLOPs. None of this includes moving stored vectors (16–32 KiB per candidate). We have not measured end-to-end latency.
    2. Counts follow Table 1 of the paper and precede compression and index duplication. ColBERT is shown with 128-d token vectors and 32-token queries; BreadBowl-Embed with 16 slots of 256 routing + 256 value dimensions. Scoring cost: ColBERT 32 × n × 128 multiply-adds; BreadBowl-Embed 16 × 16 × 256 (routing) + 16 × 16 × 256 (value read) + 16 × 256 (comparison) per passage. At 384 tokens a 2-bit ColBERTv2 index needs ≈13.5 KiB per passage, versus 16 KiB for BreadBowl-Embed's slots in bf16.
    3. Demo numbers come from running the released weights through a NumPy re-implementation of the inference path, in float32 on CPU. Its tokenization matches the reference token IDs, and its routing vectors agree with the reference GPU outputs (bf16) at a per-slot cosine of 0.9998 or better, so values can differ from the paper's figures in the third decimal.
    4. Official BEIR full corpora and relevance judgments, test splits except MS MARCO (dev). nDCG computed with pytrec_eval. CQADupStack's 12 forums are averaged into one task before the 15-task macro average. Candidates come from Algorithm 1 of the paper (128 slot hits per query slot, a 1,024-document shortlist, the top 100 by routing score) with exact flat search. One fixed checkpoint and one set of retrieval settings for every task; results are point estimates.
    5. 5,000 development queries from nine sources (LightOn subsets of MS MARCO, NQ, TriviaQA, SQuAD v2, HotpotQA and FEVER, plus StackExchange duplicates, SciRepEval search and PubMedQA), each searched over its own prepared corpus (3.41M documents in total). They were held out before candidate mining but come from the same sources as the training data, and they were used for model selection. The top-1 fix/break counts are our own tally from the same runs.
    6. The checkpoint is published as breadbowl-embed-v1.2-preview. "v1.2" is the internal research iteration; this is BreadBowl's first public model release.