Case study, 2026
Evalgate
A Rust evaluation runner and CI gate for live or recorded model outputs. Define weighted checks in YAML, classify repeated runs, and report score regressions.
01
The problem
Prompt changes ship without tests. A wording tweak in the system message, a new few-shot example, or a model swap can pass code review and break production behavior. Regressions show up as customer complaints, not CI failures.
Evaluation should support live endpoints during development without making CI depend on API keys. Evalgate uses the same YAML schema for OpenAI-compatible, Anthropic, echo, and recorded replay targets. The replay path keeps pull request runs deterministic.
02
How it is built
The workspace separates pure evaluation logic, asynchronous IO, and browser bindings:
- evalgate-core. Validates suites, renders templates, evaluates checks, scores weighted results, classifies repeats, and builds reports without doing IO.
- evalgate-cli. Executes live or replay targets with bounded concurrency, rate limiting, retries, request timeouts, disk caching, and snapshot storage.
- evalgate-wasm. Exposes deterministic evaluation, report comparison, and a flaky classification example to the browser demo.
Decisions that mattered
-
Repeat classification.
--repeatrecords a case pass rate and distinguishes stable passes, stable failures, and partial flaky results. - Checks at different layers. Text, regex, JSON shape, similarity, latency, cost, snapshot, LLM judge, and embedding checks feed one weighted score.
- CI output. Runs render terminal tables, JSON, or JUnit XML. Comparisons render tables, JSON, or Markdown and exit with status 1 on regression.
03
What was hard
Parts that took more than one pass.
Keeping network behavior testable
Target requests share a concurrency semaphore and token interval, while retry policy handles HTTP 429 and server errors. Local HTTP tests verify request spacing, retry order, provider response parsing, and cache hits.
Flaky cases need their own status
A case that passes some samples is different from a stable regression. The report keeps pass rate and flaky state alongside weighted check results, so comparison output does not collapse intermittent behavior into one verdict.
Snapshots are explicit
Snapshot checks accept a first output only with --update-snapshots,
then compare later output against that saved value.
04, Numbers
Measured, then stated.
1.387 ms
to evaluate 10,000 deterministic checks natively (release build)
53
passing tests, including 8 local HTTP integration tests
1,073,420 bytes
optimized WASM binary with the size-focused release profile
15
named check kinds in the suite schema