The Boring Parts / Catch Engineering
Modeling

Owning the second opinion

Why the models that check our assistant's messages are trained by us and run on our own machines, and what it takes to answer one without writing a word

Dor Bernsohn · X · 8 min read
XLinkedIn

Every message our assistant sends is checked first, by a model we trained and run ourselves. It answers in about half a second, two to three times quicker than the cheapest hosted model we could call for the job and roughly seven times quicker than the reasoning ones, on hardware where the conversation never leaves our infrastructure.

An AI assistant that writes to other people on your behalf will, sooner or later, get something wrong. At Catch, the assistant writes for executives, to their investors, their board and their customers, so a wrong message is not a small thing. Here is one it nearly sent:

“Thursday at 3 works for Dana — I’ve put it in her calendar, see you then.”

Fluent, plausible, and nobody had opened Dana’s calendar. Sent, it books a meeting that may not exist in the other person’s week, and the executive finds out when someone doesn’t turn up. The check that asks “were these meeting times actually verified?” objected, the message was never sent, and the assistant rewrote it to propose the time rather than confirm it.

A check catching a time the assistant never verifiedThe assistant drafts a reply saying Thursday at 3 works. Four checks run in parallel; the one asking whether the meeting times were actually verified objects, because the executive's calendar was never read. The reply is not sent. The assistant rewrites it to propose the time instead of confirming it, the checks pass, and that version goes out.DRAFT“Thursday at 3 works for Dana — I’ve put it in hercalendar, see you then.”CHECKS, IN PARALLELNames real?Times verified?objectsRight recipient?Promise recorded?Nobody opened Dana’s calendar. The reply states atime as agreed that was never checked — not sent.REWRITTEN, THEN SENT“Dana has Thursday at 3 open — shall I book it andsend the invite?”All checks passon the second draft
Fig. 1The assistant drafts a reply confirming a time nobody verified. Four checks run in parallel; the one asking whether the times were verified objects, so the reply is rewritten to propose the time instead, and that version passes.

In the previous post, Nate described how we limit what an agent is allowed to see. This post is about the other half: checking what it says before it goes out.

Every reply passes through a set of these checks before it is sent. Each one asks a single yes-or-no question. Were these meeting times actually verified? Does every name in this message appear in the conversation? If any check objects, the message is not sent and the assistant rewrites it. They run in parallel, so a reply waits for the slowest check rather than for all of them.

The obvious way to do this

Send each question to the best model you can buy and let it judge. That is where we started, and for a prototype it works well.

Why it doesn’t hold

A check is not a normal model call. Four things it needs are hard to get from a hosted model:

  • Fast. A check sits between the assistant and the send button, and since the checks run in parallel, the slowest one decides how long the user waits.
  • Accurate. A check that raises false alarms gets ignored or switched off, which is worse than not having it.
  • Consistent. The same message should get the same answer next week. A provider can change the weights, the safety tuning or the serving stack and keep the model’s name; from the outside you see only the effects. Studies have measured that drift over a few months, and with no change at all, one hosted reasoning model we used answered 10% to 20% of borderline cases differently on a re-run.
  • Private. To know whether a name in a reply is real, a check has to read the whole conversation, not just the reply. That makes it one of the most privileged readers in the system.

The cheap shape everyone is reaching for

Checks are a lopsided problem. Over a recent week our model answered about 190,000 of them, most alongside a hosted checker whose verdict was the one that counted, and roughly 1.8% of its answers objected. The other 98% wrote out reasoning nobody needed.

So the interest right now is in not writing it: attach a scoring layer to the model and get one pass over the conversation, a number between zero and one, nothing else. TypeSafe calls this a System One model, after the fast, intuitive half of human judgement, and their Jev model is the version we tested against.

So we measured Jev on our own checks: the same held-out split we score our own models on, 4,939 of them, 476 warns, asked as one yes-or-no question each. It is as fast as advertised, and faster than our own model at the tail. It is also not close on the judgement.

Our 27B headJev
Traffic cleared at ≤1% warn-miss64.9%19.7%
Traffic cleared at ≤0.5%57.4%15.8%
AUROC0.9840.872
Best warn F10.8100.539
Median answer201 ms323 ms
90th percentile694 ms410 ms

Ours is a forward pass on one B300, the GPU production serves on, timed one request at a time over 1,500 prompts; Jev’s is a round trip to a shared service, and it answered 99.9% of the split. The speed is real, and it is the part you can buy. What it cannot buy is a model that knows our rules: Jev has never seen them, and at a budget of missing at most one warn in a hundred it can wave through a fifth of the traffic where ours waves through two thirds.

Why we train our own

A classifier you can buy answers a question that isn’t ours. Our checks are about our rules: what counts as a time the assistant actually verified, which names had to appear in the conversation before it could use them. Those rules move as the product does. The choice is not whether the model is small, it is whose judgement it carries and how fast we can change it. So we train it on our own traffic and run it on hardware we operate: the conversation stays in our infrastructure, and the number we promote a version on is how often it agrees with the teacher on conversations it never saw. With a hosted model that number is something you can measure but not move.

Teaching it

The best model we have is the teacher, not the check. Real traffic gives us the questions, the teacher answers them, and a much smaller open-weights model is trained on those answers and on the reasoning behind them. This is distillation, and small, dedicated safety models are not new. Llama Guard is a well-known example.

Where the labels come fromReal checks are sampled, flagged messages plus a sample of the rest, answered offline by the teacher with its reasoning kept, and filtered, then split into train and test sets that share no conversation. The filter drops empty or cut-off reasoning and labels written for older instructions, and limits easy approvals.Productionflagged messages+ a sample of the restTeacheranswers offline,keeps its reasoningFiltercheck andrelabelTrain / testno sharedconversationsLabels we throw out• reasoning came back empty• reasoning was cut off• written for older instructions• too many easy "this is fine" examples
Fig. 2Real checks supply the examples, the teacher supplies the answers and reasoning, and a filter removes weak labels before training.

Getting that reasoning out of a teacher is the awkward part. OpenAI decided not to show the raw chain of thought of its reasoning models, and Anthropic’s API can return the reasoning empty, so some of what we pay for never arrives. Our own model has no such setting: we see every token it writes, on every call, which is what lets us train on it and read it back when a check goes wrong.

Running it

The checks are moving onto that model a check at a time: it answers alongside most of them today and decides the ones it has proved itself on.

It is also the fastest option we have. Over three days of production traffic a single check on it took about half a second at the median, against roughly a second for the cheapest hosted model we use and three and a half or more for the two reasoning ones. Each handles a different mix of checks, so the chart compares them at the same prompt length, with the classifier head on it too: the generative checker is two to three times faster than Flash-Lite on short and medium prompts and slightly behind at the longest, and the head is quicker than both at every length.

Latency by prompt lengthMedian latency of a single check, in seconds, by prompt length in tokens. The served models are three days of production traffic; the classifier head is benchmarked on one B300. Ours: generative checker 1–2k 0.28s, 2–4k 0.39s, 4–8k 0.50s, 8–16k 0.77s, 16–32k 1.27s; Ours: classifier head 1–2k 0.13s, 2–4k 0.16s, 4–8k 0.31s, 8–16k 0.64s; Gemini 3.1 Flash-Lite 1–2k 0.95s, 2–4k 0.92s, 4–8k 0.99s, 8–16k 1.03s, 16–32k 1.16s, 32k+ 1.33s; GPT-5.6 Luna 2–4k 2.83s, 4–8k 3.09s, 8–16k 3.65s, 16–32k 4.04s; Claude Haiku 4.5 2–4k 3.29s, 4–8k 3.80s, 8–16k 4.89s. Bands with fewer than 100 calls are omitted.0s1s2s3s4s5s6s1–2k2–4k4–8k8–16k16–32k32k+Prompt length (tokens)Median latencyClaude Haiku 4.5GPT-5.6 LunaGemini Flash-LiteOurs: generative checkerOurs: classifier headGemini 3.1 Flash-LiteGPT-5.6 LunaClaude Haiku 4.5
Fig. 3Median latency of a single check by prompt length, for both of ours and the three hosted models. The generative checker and the hosted models are three days of live traffic; the classifier head is benchmarked on one B300, the GPU production serves on. Bands with fewer than 100 calls are left out.
Two things trained on one studentThe fine-tuned student checker is trained once. A classifier head trained on top of it answers a check in one pass without writing anything, and a speculative-decoding drafter trained on the student's own outputs makes the written verdict faster.Studentour fine-tuned checker,distilled from the teachertrained on toptrained on its outputsClassifier headanswers a check without writing anythingDraftermakes the written verdict faster
Fig. 4One student, two things trained on top of it: a classifier head that answers a check without writing anything, and a drafter that makes the written verdict faster.

Some of that speed is a second model. A check spends most of its time writing its verdict one token at a time, so a much smaller drafter runs beside it, guesses the next few, and the checker verifies them in a single pass. We train the drafters on the checker’s own outputs, two kinds of them to see which guessed better: an EAGLE 3.1 head and a DFlash drafter. The winner lands about three tokens a pass, so the same verdict takes a third of the steps.

One rule holds throughout: when in doubt, go to the stronger model, never guess. A question like “does this name appear in the conversation?” is decided against the conversation itself, and that evidence is never trimmed to fit a smaller window. Dropping half of it would not make the answer slightly worse, it would invert it: every name in the half we cut reads as invented. When it doesn’t fit, the check goes to the larger model.

Making the common case free

This is where the scoring layer comes back, in front of the checks rather than instead of them. A classifier head on top of our fine-tuned 27B model scores a short reply in 135 milliseconds and a long one in 642, reading the whole conversation to do it, and writing nothing.

It does not have to be right on its own, only confident about the easy cases, which is nearly all of them. Ranking held-out conversations by its score, we can clear about two thirds of all checks and still catch 99 of every 100 objections the full check would have raised. The rest, the replies it thinks are wrong and the ones it is unsure about, go to the generative model, which is the only thing that can say why a message is wrong.

A fast router in front of the checkA classification head on our trained model scores the drafted reply in about under 400 ms. Confidently fine replies are cleared without generating anything; objections and uncertain scores are escalated to the full check, which writes the reason.Drafted replyplus the conversationScoring headon our 27B model~200ms, no textconfidently fineobjection or uncertainClearedabout two thirdsFull checkwrites the reason
Fig. 5A scoring layer on our trained model reads the drafted reply and the conversation in under 400 milliseconds. Confidently fine replies are cleared; objections and uncertain scores go to the full check, which writes the reason.

A skipped check is a message nobody looked at for that one question, so how many objections we are willing to miss is set per question rather than once. What that buys varies more than we expected: at the same budget it clears nearly 90% of the traffic on the easiest question and about 17% on the hardest.

The cost

None of this is free. It takes a labelling pipeline, GPUs to keep running, and a retraining cycle whenever the rules change, and running models ourselves doesn’t make them perfectly repeatable either, since inference has its own sources of randomness. What we get back is control. When a check behaves differently, we can nearly always trace it to something we changed, and see what. For an assistant that speaks on behalf of executives, we think that is worth it.

Not using Catch yet?

References

  1. Studies have measured arxiv.org
  2. calls this a System One model typesafe.ai
  3. distillation arxiv.org
  4. Llama Guard arxiv.org
  5. decided not to show simonwillison.net
  6. can return the reasoning empty platform.claude.com
  7. EAGLE 3.1 vllm.ai
  8. DFlash arxiv.org
  9. inference has its own sources of randomness thinkingmachines.ai