Case study, 2026
Screener
A hybrid search engine in Rust: BM25 with typo tolerance and ranking rules on the keyword side, a static embedding model on the semantic side, fused with reciprocal rank fusion. One core library, a CLI, and a WASM build that runs the demo in the browser without a server round trip.
01
The problem
Keyword search misses paraphrases. "Policy gradient training" will not match a paper that only says "reinforcement learning" in the abstract. Vector search fixes that, but it misses exact terms and typos. A query for "informtion retrieval" should still find information retrieval papers.
Most products need both signals. The hard part is not calling two APIs. It is ranking: when to trust BM25, when to trust cosine similarity, and how to break ties without hiding why a result appeared.
02
Design
The pipeline splits cleanly into build time and query time.
- Index build. Tokenize titles and bodies into an inverted BM25 index. Embed each document with a static model (HuggingFace tokenizer, int8 lookup table, mean pool, L2 normalize). Serialize document embeddings to a postcard blob; ship the model as sidecar files.
- Lexical retrieval. Parse the query into terms. Match the dictionary with an FST plus Levenshtein automaton (bounded distance by term length; first-letter mismatches cost double). Prefix search on the last term. Score with BM25, then bucket-sort by words, typo, proximity, attribute rank, word position, exactness.
- Vector retrieval. Brute-force cosine similarity over document embeddings. Fine at 3,000 docs; HNSW deferred.
- Fusion. Reciprocal rank fusion with k=60 and a tunable semantic ratio. Each hit carries bm25 rank, vector rank, fused score, and per-term typo counts.
screener-core has no IO. The CLI builds indexes and runs evals.
screener-wasm exposes one constructor and a query method
that returns JSON.
03
What was hard
Details the README skips.
WordPiece on full abstracts
The first index build hung after "Indexed 3000 documents." Greedy WordPiece over multi-kilobyte abstracts is quadratic in text length. Truncating embed input to 512 characters fixed build time without hurting recall much on titles plus lead sentences.
Tokenizer parity with the reference
Semantic recall near zero turned out to be a bug, not a weak model. A hand-rolled WordPiece tokenizer did not match the reference encoder. Switching to the HuggingFace tokenizer crate and excluding special tokens from the mean brought cosine parity above 0.99 on a checked-in fixture.
FST build on load
Building the term FST eagerly during engine construction trapped WASM with an unreachable panic on the full corpus. Lazy initialization on first query fixed it without changing search behavior.
04, Numbers
Measured on the eval set.
Known-item search over 300 queries (100 sampled docs, three query types each: keyword, typo, paraphrase). Hybrid semantic ratio tuned to 0.3 on the train split; numbers below are on the held-out test split (150 queries).
| Query type | Mode | Recall@1 | Recall@10 | MRR |
|---|---|---|---|---|
| Keyword | Keyword | 0.688 | 0.896 | 0.753 |
| Keyword | Semantic | 0.583 | 0.854 | 0.706 |
| Keyword | Hybrid, rank fusion | 0.792 | 0.938 | 0.852 |
| Keyword | Hybrid, score blend | 0.750 | 0.958 | 0.822 |
| Typo | Keyword | 0.646 | 0.958 | 0.749 |
| Typo | Semantic | 0.396 | 0.812 | 0.542 |
| Typo | Hybrid, rank fusion | 0.750 | 0.958 | 0.829 |
| Typo | Hybrid, score blend | 0.771 | 0.938 | 0.842 |
| Paraphrase | Keyword | 0.625 | 0.896 | 0.734 |
| Paraphrase | Semantic | 0.729 | 0.917 | 0.793 |
| Paraphrase | Hybrid, rank fusion | 0.750 | 0.917 | 0.816 |
| Paraphrase | Hybrid, score blend | 0.792 | 0.917 | 0.844 |
The train split selected a 0.3 semantic mix with train MRR 0.873. The reference embedding parity test passed with cosine above 0.99 for every fixture.
10.35 MB
gzip transfer payload for the search assets
12
tests: typo rules, ranking, BM25, RRF, embed parity, round trip
9.81 MB
rebuilt postcard index
3,000
arXiv papers (cs.CL, cs.LG, cs.IR metadata)
05
Try it
Try the live demo or browse the source on GitHub.