You can't ask a model if it's sure
A model misread one character and scored itself 0.97 confident, three times running. Why a confidence gate structurally cannot catch that, and what a second model does and doesn't buy you.
A model read one character wrong on a scanned document. A digit, rendered in a font where it's nearly identical to a letter, separated only by a small feature that survives or doesn't depending on how the page was scaled on its way in. The model picked the letter.
Then it scored itself 0.97 confident. I ran it twice more and got 0.97 and 0.98, on the same wrong value, byte-identical each time.
That self-reported number was the gate deciding whether a human ever looked at the document.
(Two honesty notes up front, because both are load-bearing. This is one document, one glyph, three repeats, observed in my own setup. A real receipt, not a benchmark, and I'll flag the sample size again where it matters. And the whole system is greenfield. Nothing here is deployed, the second-model guard I describe below has never executed outside its unit tests, and no real document has been billed against any of it. Present tense below describes a design, not a running system.)
Why can't a confidence score catch a confident error?
Because calibration is a property of a distribution and your risk is a property of a document.
The intuitive complaint is that models are overconfident liars. That's the wrong critique, and at the frontier it's contestable. Tian et al. found that in models tuned with reinforcement learning from human feedback (RLHF, which is how most chat-facing models are finished), the confidence a model writes out is typically better calibrated than its own internal token probabilities (EMNLP 2023, a peer-reviewed natural-language-processing conference). Asking for a number is a legitimate thing to do. That's a 2023 result on 2023 models, and I'd expect verbalized confidence to have improved since, which makes my point harder to argue rather than easier.
The problem is what a well-calibrated number buys you. A perfectly honest gate that says 95% is telling you that across many documents it's wrong about one in twenty. It tells you nothing about this document, which is the only one the design would auto-approve and bill against. A gate can be flawlessly honest in aggregate and still hand you a wrong critical field at exactly the advertised rate.
You also can't fix it by asking harder, and the reason is structural. The error and the confidence come out of the same forward pass, so a second opinion from the same model is the first opinion wearing a hat.
The obvious citation here is subtler than it looks, so let me be precise about what it does and doesn't say.
Gorbett & Jana benchmark P(True) — Kadavath et al.'s ask-the-model-outright baseline, where you simply ask a model whether a proposed answer is correct. They score it in AUROC, area under the receiver operating characteristic curve, which measures how well a signal separates right answers from wrong ones. 1.0 is perfect. 0.5 is a coin flip.
P(True) scores 0.518 on MMLU (Massive Multitask Language Understanding, a broad general-knowledge test) across their capability-gap pairs, and a mean of 0.463 on GSM8K (grade-school maths word problems) across all pairs (arXiv:2603.25450). So it's a coin flip with better marketing on the first, and on the maths it lands below the 0.504 they measure for random guessing.
One important caveat on borrowing those numbers, though. In their setup the question goes to a second model looking at the first one's answer, not to the model grading itself. So it isn't direct evidence about self-scoring. What it is evidence for is narrower and more useful: asking a model outright "is this right?" performs near chance even when the model doing the asking is a different one. Comparing two independent reads is a different move, and it scores considerably better.
My own run says it more bluntly than the literature does. On the failure I cared about, the model asserted 0.97 or better on a wrong value, three times out of three.
What does a second model actually buy you?
An independent read. That's the whole product.
The design is deliberately dumb. Run a second, different model over the critical fields. Compare the values after normalizing format, so leading zeros, casing and separators don't count as disagreements. Boring, and load-bearing. If the two reads differ, the document stops being eligible for auto-approval and goes to a human.
The check never picks a winner. That constraint is the entire safety argument, and it's the one people are most tempted to relax. If the machine adjudicates between two candidate readings it can "correct" a genuinely right value into a wrong one, and you've built a system that introduces errors while wearing the costume of a system that catches them. Detect, don't adjudicate.
On a mismatch, the routing splits. A critical-field disagreement escalates to a human. A non-critical one ships what the first scan captured. I'll be straight about what that second branch is. The machine resolves a disagreement by fiat in favor of scan one. That's an accepted-error policy, and it isn't detect-don't-adjudicate. The honest formulation is that I adjudicate where being wrong is cheap and escalate where it isn't. Defensible. Just don't call it "never adjudicates."
There's outside support for the exact shape of this. In Gorbett & Jana's comparison, the baseline that matches my design (two full independent generations, string-compared, which they call V-Agree) scores 0.756 mean AUROC on MMLU across their capability-gap pairs, the same slice P(True) scored 0.518 on, edging out the cheaper single-pass method the paper actually proposes. The dumb version holds its own.
The caveats travel with that number. Their models are small open-weight ones, 8 billion parameters and under, well below the frontier models a design like mine would actually use. Most pairs are asymmetric weak-to-strong. Their cheaper single-pass proxy, which is the method most of the per-pair detail is reported for, runs 0.584 to 0.896 across a different slice, the same-size cross-family pairs on MMLU. Change benchmark and it can fall through the floor. One pair whose two models sat a single accuracy point apart, 64% against 65%, scored 0.421 on TriviaQA (a trivia question-answering benchmark) with context supplied, below chance, against a baseline of 0.847 for the simpler trick of watching one model's own token-level uncertainty. Which model you pair with which is most of the result.
One more gap worth naming before I lean on any of this. Every paper I cite here tests text models, mostly on question-answering benchmarks. None of them tests vision-based document extraction. The mechanisms transfer plausibly. They haven't been shown to.
Which documents get the second look?
The design fires on the one source with a track record of this failure.
That's much cheaper than gating on a confidence band, and it's honest about what I actually know. The tempting version of this sentence is a general claim about bad scans and unusual fonts clustering together. I don't have evidence for that. What I have is one source that has been problematic, and a trigger that follows the evidence I've got.
Say the limitation out loud. It buys zero coverage of the same failure at any other source.
Which raises the obvious question, and it's the one I'd ask if someone showed me this design. If exactly one source is the problem, why not send that source's critical fields to a human and skip the second model entirely?
I want to answer that honestly rather than well, because the honest answer is weaker. A second model is a filter: it escalates only the documents where two reads disagree, so you pay a person for a slice instead of for the whole source. That's the argument, and whether it holds depends entirely on how thin the slice is. If the two models disagree on 2% of that source, the filter is obviously right. If they disagree on 40%, I've added an API call and a pipeline stage to arrive at the same review queue I'd have had for free.
I don't know which it is. One problematic source could be a small enough share of volume that human review was affordable all along, and I never priced it. The filter argument is sound in shape and unquantified in fact, and if I were the person signing off on this I'd want the disagreement rate before the design, not after.
There's also a tension worth sitting with, and it's conditional on something I don't know. If that source produces failures because its inputs are genuinely more ambiguous, then pointing the guard there points it at the slice where it's least reliable. Both models get the same bad stimulus, so both have the same reason to land on the same wrong answer. The spend would be concentrated where the errors are and where the guard is weakest, and those would be the same place. I can't tell you whether that's what's happening. It's the version I'd want to rule out before trusting the guard too far.
Worth naming one more wrinkle, because it looks like a contradiction. The guard only runs on documents already heading for auto-approval, which means documents the confidence gate cleared. I've spent the opening of this post on why that gate can't be trusted. It still works as a cheap first cut. It just can't be the only one.
Why not a third model and a majority vote?
Because the vote is the problem, not the count. A majority vote is the machine adjudicating, which is the one thing this design refuses to do, and three models can agree on the same confident error.
This isn't a hunch. Kim et al. found that on one leaderboard dataset models agree 60% of the time when both are wrong, and that larger, more accurate models have highly correlated errors even with distinct architectures and providers (ICML 2025, the International Conference on Machine Learning). Goel et al. found error correlation rising as capability rises, which matters because it means the problem gets worse as the models get better (ICML 2025). A single-author preprint sharpens it further, naming the slice of inputs where the entire market fails together, a blind spot that no pairwise agreement statistic can detect (arXiv:2606.27288, unreviewed, so weigh it accordingly).
Add a third model and you've bought a tiebreaker that shares the tie. The human is the tiebreaker.
What the guard can't catch
Its coverage is bounded by the decorrelation of the pair, by how differently the two models actually fail. The count of models is not the variable, and I haven't measured the decorrelation either.
The failure I built this for is the worst possible case for the design, and I walked into it knowingly. A glyph that looked genuinely ambiguous to me in the pixels presents an identical stimulus to both models. They have a strong shared reason to make the same mistake. Hamidieh et al. found cross-model disagreement is diagnostic precisely when input uncertainty is low (arXiv:2604.17112), which points the wrong way for this design. Again, that's a text benchmark, and "ambiguous pixels" is my analogy for their aleatoric uncertainty (noise in the input itself, as opposed to the model's own ignorance) rather than their finding.
So I predicted the two models would agree on the same wrong value.
They didn't, which cuts against the worry I raised earlier, at n=1, which is to say barely. The primary reproduced the misread on 3 of 3 repeats. The cross-check read it correctly on all three. Agreement on the wrong value was zero. Under the design's rule that's a demotion to review, which is the point. The cross-check's clean 3 of 3 is also weak evidence the glyph was more legible than I'd judged it.
I don't know why one read it wrong and the other didn't. Same pixels, same prompt, different answer, three times each. That's the result. Any explanation I offered would be a story I hadn't tested, and I've been burned on precisely that before.
n = 1. One document, one glyph, one direction. I'd rather publish a narrow true result than a broad comfortable one.
And the thing I couldn't test is the thing that matters most. My two models are two tiers of one family from one vendor. That is not two lineages. Given Kim et al., a different vendor wouldn't buy independence by itself either, so "same family, different tier" is a substantially weaker independence assumption than it sounds. Testing a genuinely different lineage would have meant sending a real document to a new vendor, which is a data-governance call rather than an engineering one, so I stopped. If I go back it'll be with a synthetic stimulus and a sweep across rendering resolution to find where each model's answer flips. Better experiment, no governance problem.
Is a second scan worth it?
Yes for the critical fields, on an asymmetry argument, which is the only kind that works here. Whether the second model beats sending that one source straight to a human is the question I raised above, and I still can't answer it.
The cost is bounded. The failure isn't. A second scan adds roughly 41% to a document's extraction spend at Anthropic's published rates as of August 2026, derived against a baseline that counts both of the pipeline's primary-model calls. Against that, it buys down an unbounded, silent, high-severity failure. Severity is what makes the asymmetry hold even at a low base rate, though I can't tell you how low. A wrong critical field promoted silently doesn't announce itself. It would surface later as a dispute and a trust cost you can't price.
I should name what this cost model leaves out. It prices the API and says nothing about the human queue. I don't have a false-positive rate worth quoting. The clean evidence is three documents with no false positives on the non-critical fields, which is different evidence from the single-document experiment above and still nowhere near a rate. Escalation volume is the number a reader would actually need before building this, and I don't have it.
Takeaways
- A confidence gate can't catch a confidently wrong answer. Calibration is aggregate; your risk is per-document. You need an independent signal, not more of the same one.
- A second opinion's coverage is its decorrelation, not its count. Three models don't beat two if they fail together. If you can't say why your two models fail differently, you don't know what your check covers.
- Detect, don't adjudicate. And know which of your detections you're quietly adjudicating anyway.
FAQ
Can I use two tiers of the same model family as a consensus check?
The pair caught it on the one document I tested, running the two models by hand, but that's n=1 and I wouldn't generalize from it. Published work finds error correlation rises with capability even across different providers, so a same-family pair is a weaker independence assumption than it looks. Pair choice dominates the result.
Why not just raise the confidence threshold instead?
Because it doesn't touch this failure. The wrong value scored 0.97 and would have cleared a 0.95 bar, a 0.96 bar, and any bar you can set without sending most of your clean documents to a human too. Raising the threshold trades volume for safety on the errors the model is unsure about. Confident errors sail through at every setting.