Case study 01, 2026

Flamide

An AI agent that answers student enquiries for education consultancies over Instagram, Messenger, and WhatsApp, in Nepali, English, and the romanized mix people actually type. It qualifies the lead, fills the CRM, follows up when the student goes quiet, and hands off to a counselor at the moments a machine shouldn't decide.

Role
Lead engineer. I designed and built the core architecture.
Stack
Next.js 16, TypeScript, Supabase (Postgres, RLS, Realtime), DeepSeek, Vitest, Playwright
Scale of the codebase
562 commits, 1,592 unit tests, 315 end-to-end tests
Code
Private. Happy to walk through it on a call.

01

The problem

Consultancies in Nepal run on DMs. Most enquiries come in after hours, in mixed script, and ask the same handful of questions about courses, fees, and intakes. A generic chatbot handles those fine.

The expensive messages are the other ones. "Will my visa get approved?" "I was refused last year." "Can you lower the fee?" A bot that answers those confidently is worse than no bot, because a wrong answer about a visa file has real consequences for a student and real liability for the consultancy.

So the job was never "answer everything." It was: answer what is grounded in the consultancy's own data, and escalate the rest with a reason a counselor can act on.

02

How it's built

Every AI turn goes through the same fixed order:

  1. Inbound guardrails. Prior refusals, complaints, finance questions, and requests for a human escalate before the model is called.
  2. Generate. A six-layer prompt: style, guardrails, domain reference, consultancy identity, catalog, then conversation state.
  3. Outbound guardrails. Code checks the reply for visa predictions, invented requirements, and misattributed ones.
  4. Regenerate once. One retry with the violation named. No loops.
  5. Escalate. If it still fails, the conversation goes to a human with the reason attached.

Decisions that mattered

  • Tenant isolation lives in the database. Every table carries workspace_id with a row level security policy, and composite (workspace_id, id) foreign keys make a cross-tenant reference impossible to write. The isolation suite is the most important test suite in the repo.
  • The model is not a privileged writer. Its writes go through narrow SQL functions. Internal columns are kept out of prompts by an allowlist, and a schema test fails the build if someone adds a column without classifying it.
  • The prompt is shaped for the cache. The first five layers are byte-stable per workspace and sit before a cache boundary, so the provider's prompt-cache discount applies to most of every request.
  • Scheduled delivery is a table, not a timer. Replies are claimed with a lease and sent exactly once, so typing delays, follow-up nudges, and multi-part messages survive a worker restart.
  • Channels sit behind an adapter. Nothing above ChannelAdapter knows a Meta payload exists, so every flow is testable with a mock channel before platform approval.

03

What broke, and what I changed

The parts a demo never shows.

Incident

Green suite, silent product

The first live call to the model returned nothing. Every reply, every time. In the "careful" effort mode the model spent the whole token budget reasoning and got cut off before writing a character, and the engine quietly logged that as "silent." 1,487 tests passed, because the mock provider had no token ceiling.

I sized the ceilings per effort level from measured completions (they had reached 96.8% of the old limit), added a guard that fails loudly on an empty completion, and started treating live runs as a separate gate.

Incident

Every fact true, the sentence false

A reply told a student that CQUniversity needs 75% for CBSE students. CQUniversity had no CBSE requirement on file, and 75% belonged to Coventry, in another country. Each piece was real, so an existence check could never catch it.

The guardrail now validates the whole tuple (university, board, value) against the database, and a missing row counts as a violation. That turns "requirement unknown" into an escalation instead of a guess.

Incident

The provider answered with 86 spaces

Certain JSON-mode requests came back as pure whitespace with a normal stop reason, deterministically, at temperature 0. I isolated the trigger to the JSON-mode flag with five variations of the captured payload.

The engine now retries that one turn without the flag and parses the result through the same reader. It's documented as a mitigation rather than a fix, because the provider behavior is still unexplained.

04, Numbers

Measured, then decided.

1,592

unit tests across 89 files, plus 315 end-to-end and 4 realtime

~12x

more output tokens in careful mode than fast, for no grounding gain

~13s

per careful reply, too slow for DMs, so fast became the default

5 min

follow-up floor, enforced at the constraint, validator, and UI

05

Limits, stated plainly

  • Instagram and WhatsApp adapters are stubs until Meta's app review clears. Everything runs on the mock channel today.
  • There is no live-model canary set yet. It's the next thing I'd build, since two of the worst bugs only showed up live.
  • The formal six-conversation evaluation against the current version hasn't run. The founder live session has.
Next case study Sekkin →