Case study, 2026

Evalgate

A Rust evaluation runner and CI gate for live or recorded model outputs. Define weighted checks in YAML, classify repeated runs, and report score regressions.

Role
Solo project. Design, implementation, and site integration.
Stack
Rust, Tokio, reqwest, serde_yaml, wasm-bindgen, GitHub Actions
Links
Live demo, GitHub
Code
Open source under MIT.

01

The problem

Prompt changes ship without tests. A wording tweak in the system message, a new few-shot example, or a model swap can pass code review and break production behavior. Regressions show up as customer complaints, not CI failures.

Evaluation should support live endpoints during development without making CI depend on API keys. Evalgate uses the same YAML schema for OpenAI-compatible, Anthropic, echo, and recorded replay targets. The replay path keeps pull request runs deterministic.

02

How it is built

The workspace separates pure evaluation logic, asynchronous IO, and browser bindings:

  1. evalgate-core. Validates suites, renders templates, evaluates checks, scores weighted results, classifies repeats, and builds reports without doing IO.
  2. evalgate-cli. Executes live or replay targets with bounded concurrency, rate limiting, retries, request timeouts, disk caching, and snapshot storage.
  3. evalgate-wasm. Exposes deterministic evaluation, report comparison, and a flaky classification example to the browser demo.

Decisions that mattered

  • Repeat classification. --repeat records a case pass rate and distinguishes stable passes, stable failures, and partial flaky results.
  • Checks at different layers. Text, regex, JSON shape, similarity, latency, cost, snapshot, LLM judge, and embedding checks feed one weighted score.
  • CI output. Runs render terminal tables, JSON, or JUnit XML. Comparisons render tables, JSON, or Markdown and exit with status 1 on regression.

03

What was hard

Parts that took more than one pass.

Detail

Keeping network behavior testable

Target requests share a concurrency semaphore and token interval, while retry policy handles HTTP 429 and server errors. Local HTTP tests verify request spacing, retry order, provider response parsing, and cache hits.

Detail

Flaky cases need their own status

A case that passes some samples is different from a stable regression. The report keeps pass rate and flaky state alongside weighted check results, so comparison output does not collapse intermittent behavior into one verdict.

Detail

Snapshots are explicit

Snapshot checks accept a first output only with --update-snapshots, then compare later output against that saved value.

04, Numbers

Measured, then stated.

1.387 ms

to evaluate 10,000 deterministic checks natively (release build)

53

passing tests, including 8 local HTTP integration tests

1,073,420 bytes

optimized WASM binary with the size-focused release profile

15

named check kinds in the suite schema

Next case study Flamide →