Owning the second opinion
Why the models that check our assistant's messages are trained by us and run on our own machines, and what it takes to answer one without writing a word
Every message our assistant sends is checked first, by a model we trained and run ourselves. It answers in about half a second, two to three times quicker than the cheapest hosted model we could call for the job and roughly seven times quicker than the reasoning ones, on hardware where the conversation never leaves our infrastructure.
An AI assistant that writes to other people on your behalf will, sooner or later, get something wrong. At Catch, the assistant writes for executives, to their investors, their board and their customers, so a wrong message is not a small thing. Here is one it nearly sent:
“Thursday at 3 works for Dana — I’ve put it in her calendar, see you then.”
Fluent, plausible, and nobody had opened Dana’s calendar. Sent, it books a meeting that may not exist in the other person’s week, and the executive finds out when someone doesn’t turn up. The check that asks “were these meeting times actually verified?” objected, the message was never sent, and the assistant rewrote it to propose the time rather than confirm it.
In the previous post, Nate described how we limit what an agent is allowed to see. This post is about the other half: checking what it says before it goes out.
Every reply passes through a set of these checks before it is sent. Each one asks a single yes-or-no question. Were these meeting times actually verified? Does every name in this message appear in the conversation? If any check objects, the message is not sent and the assistant rewrites it. They run in parallel, so a reply waits for the slowest check rather than for all of them.
The obvious way to do this
Send each question to the best model you can buy and let it judge. That is where we started, and for a prototype it works well.
Why it doesn’t hold
A check is not a normal model call. Four things it needs are hard to get from a hosted model:
- Fast. A check sits between the assistant and the send button, and since the checks run in parallel, the slowest one decides how long the user waits.
- Accurate. A check that raises false alarms gets ignored or switched off, which is worse than not having it.
- Consistent. The same message should get the same answer next week. A provider can change the weights, the safety tuning or the serving stack and keep the model’s name; from the outside you see only the effects. Studies have measured that drift over a few months, and with no change at all, one hosted reasoning model we used answered 10% to 20% of borderline cases differently on a re-run.
- Private. To know whether a name in a reply is real, a check has to read the whole conversation, not just the reply. That makes it one of the most privileged readers in the system.
The cheap shape everyone is reaching for
Checks are a lopsided problem. Over a recent week our model answered about 190,000 of them, most alongside a hosted checker whose verdict was the one that counted, and roughly 1.8% of its answers objected. The other 98% wrote out reasoning nobody needed.
So the interest right now is in not writing it: attach a scoring layer to the model and get one pass over the conversation, a number between zero and one, nothing else. TypeSafe calls this a System One model, after the fast, intuitive half of human judgement, and their Jev model is the version we tested against.
So we measured Jev on our own checks: the same held-out split we score our own models on, 4,939 of them, 476 warns, asked as one yes-or-no question each. It is as fast as advertised, and faster than our own model at the tail. It is also not close on the judgement.
| Our 27B head | Jev | |
|---|---|---|
| Traffic cleared at ≤1% warn-miss | 64.9% | 19.7% |
| Traffic cleared at ≤0.5% | 57.4% | 15.8% |
| AUROC | 0.984 | 0.872 |
| Best warn F1 | 0.810 | 0.539 |
| Median answer | 201 ms | 323 ms |
| 90th percentile | 694 ms | 410 ms |
Ours is a forward pass on one B300, the GPU production serves on, timed one request at a time over 1,500 prompts; Jev’s is a round trip to a shared service, and it answered 99.9% of the split. The speed is real, and it is the part you can buy. What it cannot buy is a model that knows our rules: Jev has never seen them, and at a budget of missing at most one warn in a hundred it can wave through a fifth of the traffic where ours waves through two thirds.
Why we train our own
A classifier you can buy answers a question that isn’t ours. Our checks are about our rules: what counts as a time the assistant actually verified, which names had to appear in the conversation before it could use them. Those rules move as the product does. The choice is not whether the model is small, it is whose judgement it carries and how fast we can change it. So we train it on our own traffic and run it on hardware we operate: the conversation stays in our infrastructure, and the number we promote a version on is how often it agrees with the teacher on conversations it never saw. With a hosted model that number is something you can measure but not move.
Teaching it
The best model we have is the teacher, not the check. Real traffic gives us the questions, the teacher answers them, and a much smaller open-weights model is trained on those answers and on the reasoning behind them. This is distillation, and small, dedicated safety models are not new. Llama Guard is a well-known example.
Getting that reasoning out of a teacher is the awkward part. OpenAI decided not to show the raw chain of thought of its reasoning models, and Anthropic’s API can return the reasoning empty, so some of what we pay for never arrives. Our own model has no such setting: we see every token it writes, on every call, which is what lets us train on it and read it back when a check goes wrong.
Running it
The checks are moving onto that model a check at a time: it answers alongside most of them today and decides the ones it has proved itself on.
It is also the fastest option we have. Over three days of production traffic a single check on it took about half a second at the median, against roughly a second for the cheapest hosted model we use and three and a half or more for the two reasoning ones. Each handles a different mix of checks, so the chart compares them at the same prompt length, with the classifier head on it too: the generative checker is two to three times faster than Flash-Lite on short and medium prompts and slightly behind at the longest, and the head is quicker than both at every length.
Some of that speed is a second model. A check spends most of its time writing its verdict one token at a time, so a much smaller drafter runs beside it, guesses the next few, and the checker verifies them in a single pass. We train the drafters on the checker’s own outputs, two kinds of them to see which guessed better: an EAGLE 3.1 head and a DFlash drafter. The winner lands about three tokens a pass, so the same verdict takes a third of the steps.
One rule holds throughout: when in doubt, go to the stronger model, never guess. A question like “does this name appear in the conversation?” is decided against the conversation itself, and that evidence is never trimmed to fit a smaller window. Dropping half of it would not make the answer slightly worse, it would invert it: every name in the half we cut reads as invented. When it doesn’t fit, the check goes to the larger model.
Making the common case free
This is where the scoring layer comes back, in front of the checks rather than instead of them. A classifier head on top of our fine-tuned 27B model scores a short reply in 135 milliseconds and a long one in 642, reading the whole conversation to do it, and writing nothing.
It does not have to be right on its own, only confident about the easy cases, which is nearly all of them. Ranking held-out conversations by its score, we can clear about two thirds of all checks and still catch 99 of every 100 objections the full check would have raised. The rest, the replies it thinks are wrong and the ones it is unsure about, go to the generative model, which is the only thing that can say why a message is wrong.
A skipped check is a message nobody looked at for that one question, so how many objections we are willing to miss is set per question rather than once. What that buys varies more than we expected: at the same budget it clears nearly 90% of the traffic on the easiest question and about 17% on the hardest.
The cost
None of this is free. It takes a labelling pipeline, GPUs to keep running, and a retraining cycle whenever the rules change, and running models ourselves doesn’t make them perfectly repeatable either, since inference has its own sources of randomness. What we get back is control. When a check behaves differently, we can nearly always trace it to something we changed, and see what. For an assistant that speaks on behalf of executives, we think that is worth it.
References
- Studies have measured arxiv.org
- calls this a System One model typesafe.ai
- distillation arxiv.org
- Llama Guard arxiv.org
- decided not to show simonwillison.net
- can return the reasoning empty platform.claude.com
- EAGLE 3.1 vllm.ai
- DFlash arxiv.org
- inference has its own sources of randomness thinkingmachines.ai