
We ship the kind of AI we harden.
Tavern Scribe is our own voice-first AI platform for tabletop RPGs. We run AI orchestration, realtime voice, and RAG in production every day. And because we are the client on this one, nothing below waits on approval: these are our production numbers, published in full.
Frontier model quality at a bill that broke consumer pricing.
Tavern Scribe turns multi-hour conversational audio into structured records: the events that happened, the entities involved, what changed, what was said. One recording becomes a transcript on the order of 165,000 tokens, and the pipeline runs dozens of model calls over it.
Frontier models are genuinely excellent at that work. The problem is arithmetic: at roughly $6.45 of model spend per session, gross margin sat near 50% before paying for anything else, and every new customer made the bill bigger at the same rate it made revenue bigger. That breaks consumer pricing.

Cheap model finds, expensive model thinks.
We did not replace the frontier models. We stopped feeding them haystacks: small models find the evidence, the expensive model does the thinking, and a fallback ladder keeps quality floored at the system we already trusted.
The platform
A voice-first AI platform for tabletop RPGs: AI orchestration, realtime voice, and RAG, running in production every day. We ship the kind of AI we harden for clients.
Small task-specific models
Span-pointing encoders around 200M parameters, served on CPU, trained on data we had already paid frontier models to produce. They point at evidence in the transcript; a model that can only point at your document cannot make things up.
Never-degrade fallbacks
Any small-model failure falls back to the frontier path automatically: a timeout, a failed validation, an empty serve. The worst case is the old cost for one step, never the old quality.
Shadow metrics and replay tooling
Every small model ran in shadow against live frontier calls before serving, and an isolated replay harness tests the full serve path against recorded production inputs in about 2 minutes instead of a multi-hour pipeline run.
Where the engineering effort goes.
Measured in production, published in full.
These come from real customer workloads, not benchmarks, and the one task model that did not clear the quality gate stays on the frontier path by design. The full engineering story, including the failures, is written up in the article linked below.
$6.45 to $1.60
AI processing cost per multi-hour session
~50% to ~88%
Projected gross margin, checked against measured reality
6 of 7
Task-specific models above .90 recall against human-validated reference sets
208 vs 21
Candidate events our small span-pointing model surfaced vs a budget LLM on the same 3.2-hour recording

Unit economics from the engineering write-up. Left: cost per multi-hour session, before and after. Right: projected vs measured on the cheapest tier.
Want your pipeline's version of $6.45 to $1.60?
If you run LLMs over long documents, recordings, tickets, contracts, or logs, this pattern probably applies to you. Engagements start at $4k, most full builds land near $40k, and the number we quote is the number you pay.