adding a reranker is an easy yes in a meeting. you pick a query where the right paragraph sits at rank eight, score the top 40 with a cross-encoder, and the paragraph moves to rank one, which is enough for the meeting to mark retrieval solved. on real traffic the assistant waits longer before the first token, the search portion of the bill starts to resemble the generation portion, and many of the queries people complain about look the same as they did on friday. i have shipped that path, and it held up on the queries picked for the meeting.
the second pass and its bill
a first-stage retriever, keyword or vector or the plain combination of the two, pulls a shortlist out of a large index. the reranker scores each passage in that shortlist with the query in one forward pass. a bi-encoder embeds the query and the passage separately, so it can run across a big index and still rank near misses highly. a cross-encoder sees the pair together, so it can tell that a passage shares the words and still answers a different question. an instruction-tuned model asked for a relevance score can imitate that job, usually for more money, because you are renting a general model to emit a number.
you run the reranker on 20 or 50 candidates. each candidate is a few hundred tokens, and on a hosted rerank endpoint the request is often fatter than the answer you generate afterward. you pay to read the pile, and the answer model reads the winners again.
i timed this on an internal assistant that already had hybrid retrieval and a filter for published docs. the shortlist came back in about 70 milliseconds. reranking the top 40 chunks added a little over 800 milliseconds, and those 800 sat in front of the answer model, which still needed about a second before it wrote anything. time to first token includes the whole rerank, since no answer token exists until the reranker returns. a person waiting on the reply could feel that pause. they had no view of whether citation three had moved.
the invoice followed the trace. retrieval was close to a rounding error. those 40 passages, at roughly 280 tokens each, were the line that moved. on most queries the shortlist already held a passage worth citing somewhere in the top five, which is the standard a person uses when they are not opening a ticket. the reranker was being paid to reshuffle lists that had already cleared that standard.
running the reranker on your own machines moves the spend onto hardware you keep warm. a gpu waiting for bursty questions is idle between bursts, and it is a queue when the questions land together. i have returned an error to a user because the rerank call exceeded its timeout, with a usable shortlist already in the log. retrieval had finished, and the answer never ran.
there is a second copy of the bill in the prompt. if you rerank 50 chunks and then put 20 of them into the answer model so nothing gets missed, you paid for an ordering and then kept most of the pile anyway. generation stays expensive because the context is still large. when the top of the reranked list is the part you trust, send a few passages and leave the rest out.
grading only the shortlist
the comparison that makes a reranker look required is easy to build without noticing. you keep the queries where the gold passage was already retrieved, because otherwise there is nothing to promote, and you measure whether gold reached rank one. the reranker wins on that slice, since promotion is all it can do.
complaints often come from queries outside the slice. the right page never made the candidate list. it lived under a title nobody types, or the status filter was still a comment in the ingest job. scoring those 40 chunks gives you a neat order of the wrong material. the top hit matches the wording more tightly, the score goes up, and the answer model cites it more readily than it would have cited a muddier miss, and the reply that comes back is wrong in cleaner prose.
near-duplicates that disagree follow the same pattern. a current refund policy and a help-center article from 2022 can both look like responses. the older article is often more direct, because nobody has edited the exceptions back in. a reranker aimed at "does this passage answer the question" will prefer the direct page. i watched one do that on a cancellation policy. the score fit the question i had typed, and that question left out the date. i was briefly pleased with the number, which is a silly reaction to a stale rule.
freeze a log sample before you change the stack, so a bad result is allowed to count. fifty or a hundred real queries will take the air out of a demo, and the sample is too small to treat as a leaderboard. score the first-stage top hit by asking whether you would hand that passage to the answer model as the citation. turn the reranker on and score the same queries. put the latency percentiles in the note, and put the cost per thousand queries next to them. if the latency column is blank, you do not know yet whether the bump was worth paying for.
on the assistant i timed, the top passage id changed on about a third of the sample. about half of those changes were another chunk of the same document, and the answer was going to come out the same. the changes worth keeping were a cluster of long questions where two procedures shared vocabulary and the first stage had picked the more popular document. that cluster was a real improvement, and it was smaller than the offline delta, because the offline queries had been chosen from cases where gold was already in the shortlist.
overlapping chunks distort the bill further. neighboring windows that repeat a paragraph get scored twice, which means one paragraph produces two charges. the top two hits on some of those queries were the same sentence with different heading crumbs, scored and billed as separate passages.
a short list and a flag
add candidates only while a person scoring the sample is still changing the citation they would accept. rerank ten, then try twenty. when extra passages stop changing that decision, stop adding them. moving to 40 or 100 because a bigger pool feels safer adds latency on every query, and a cache does not remove it from the traffic you added the reranker to catch. repeated questions hit the cache, and the long tail is the part that misses it.
a lot of the apparent gain is already gone once stale pages are out of the shortlist, which is before any reranker runs. drafts, old versions, and marketing pages make the reranker spend its calls separating them from the document you wanted. a filter on status and updated date does that split more cheaply. when i enabled the reranker and left the filter for a later sprint, i was paying a model to notice a boolean the ingest job could have stored.
some corpora are a reasonable place to leave the reranker on. the collection is small, someone maintains it, the first stage already returns the right neighborhood, and the leftover mistakes are the right document with the wrong section. policy pages with several similar rules are where i keep seeing that. the top hit improved often enough that a few hundred milliseconds was a fair trade, and the list stayed around fifteen passages, because twenty did not add another acceptable citation in the sample. i would leave that configuration running over a weekend.
if the gold passage is missing from the top 50, spend the next week on the retriever that builds the list, because the reranker can only reorder pages it was given. if the passage is already in the top few and you would cite the current top hit, more scoring adds latency and money to queries that were already acceptable. sku lookups and pasted error strings land there. keyword search already has the code, so a cross-encoder on that hit repeats a lookup that finished.
if you keep the reranker, gate it on something you can see in the log: a long question, and first-stage scores sitting close together. short lookups skip the call. leave the flag in place, and record the query, whether the top passage id changed, the milliseconds added, and whether the user rephrased and tried again. after a week, turn the reranker off when the rephrase rate has stayed put.