war-story
You can't ask a model if it's sure
A model misread one character and scored itself 0.97 confident, three times running. Why a confidence gate structurally cannot catch that, and what a second model does and doesn't buy you.
war-story
A model misread one character and scored itself 0.97 confident, three times running. Why a confidence gate structurally cannot catch that, and what a second model does and doesn't buy you.
token-efficiency
A sentence in two of my agent contracts described a vendor API accurately — and had been false for ten days. Three readings of one endpoint in three weeks, and where the changelog was actually hiding.
token-efficiency
One planning run cost 2.40M tokens and 97% of them were cache traffic. Prompt-tightening acts on the 3%. The fix is architectural, and here's the break-even that tells you when it's worth building.
token-efficiency
Same ticket, matched run: Sonnet cost $3.99, Opus $10.18. A measured receipt on when the frontier model is worth 2.5x — and when it isn't.