Case study, 2026

Screener

A hybrid search engine in Rust: BM25 with typo tolerance and ranking rules on the keyword side, a static embedding model on the semantic side, fused with reciprocal rank fusion. One core library, a CLI, and a WASM build that runs the demo in the browser without a server round trip.

Role
Sole author. Design, implementation, eval, and demo.
Stack
Rust, wasm-bindgen, postcard, arXiv metadata (CC0)
Scope
3,000 papers, 300 known-item queries, 12 tests
Links
Live demo, GitHub

01

The problem

Keyword search misses paraphrases. "Policy gradient training" will not match a paper that only says "reinforcement learning" in the abstract. Vector search fixes that, but it misses exact terms and typos. A query for "informtion retrieval" should still find information retrieval papers.

Most products need both signals. The hard part is not calling two APIs. It is ranking: when to trust BM25, when to trust cosine similarity, and how to break ties without hiding why a result appeared.

02

Design

The pipeline splits cleanly into build time and query time.

  1. Index build. Tokenize titles and bodies into an inverted BM25 index. Embed each document with a static model (HuggingFace tokenizer, int8 lookup table, mean pool, L2 normalize). Serialize document embeddings to a postcard blob; ship the model as sidecar files.
  2. Lexical retrieval. Parse the query into terms. Match the dictionary with an FST plus Levenshtein automaton (bounded distance by term length; first-letter mismatches cost double). Prefix search on the last term. Score with BM25, then bucket-sort by words, typo, proximity, attribute rank, word position, exactness.
  3. Vector retrieval. Brute-force cosine similarity over document embeddings. Fine at 3,000 docs; HNSW deferred.
  4. Fusion. Reciprocal rank fusion with k=60 and a tunable semantic ratio. Each hit carries bm25 rank, vector rank, fused score, and per-term typo counts.

screener-core has no IO. The CLI builds indexes and runs evals. screener-wasm exposes one constructor and a query method that returns JSON.

03

What was hard

Details the README skips.

Build

WordPiece on full abstracts

The first index build hung after "Indexed 3000 documents." Greedy WordPiece over multi-kilobyte abstracts is quadratic in text length. Truncating embed input to 512 characters fixed build time without hurting recall much on titles plus lead sentences.

Embedder

Tokenizer parity with the reference

Semantic recall near zero turned out to be a bug, not a weak model. A hand-rolled WordPiece tokenizer did not match the reference encoder. Switching to the HuggingFace tokenizer crate and excluding special tokens from the mean brought cosine parity above 0.99 on a checked-in fixture.

WASM

FST build on load

Building the term FST eagerly during engine construction trapped WASM with an unreachable panic on the full corpus. Lazy initialization on first query fixed it without changing search behavior.

04, Numbers

Measured on the eval set.

Known-item search over 300 queries (100 sampled docs, three query types each: keyword, typo, paraphrase). Hybrid semantic ratio tuned to 0.3 on the train split; numbers below are on the held-out test split (150 queries).

Query type Mode Recall@1 Recall@10 MRR
KeywordKeyword0.6880.8960.753
KeywordSemantic0.5830.8540.706
KeywordHybrid, rank fusion0.7920.9380.852
KeywordHybrid, score blend0.7500.9580.822
TypoKeyword0.6460.9580.749
TypoSemantic0.3960.8120.542
TypoHybrid, rank fusion0.7500.9580.829
TypoHybrid, score blend0.7710.9380.842
ParaphraseKeyword0.6250.8960.734
ParaphraseSemantic0.7290.9170.793
ParaphraseHybrid, rank fusion0.7500.9170.816
ParaphraseHybrid, score blend0.7920.9170.844

The train split selected a 0.3 semantic mix with train MRR 0.873. The reference embedding parity test passed with cosine above 0.99 for every fixture.

10.35 MB

gzip transfer payload for the search assets

12

tests: typo rules, ranking, BM25, RRF, embed parity, round trip

9.81 MB

rebuilt postcard index

3,000

arXiv papers (cs.CL, cs.LG, cs.IR metadata)

05

Try it

Try the live demo or browse the source on GitHub.

Next case study Evalgate →