seven of ten still fail. plan for that.
29.3%. that is the best score on the hardest real-repo coding benchmark right now, from the best model tested. which means the same model fails 70.7% of the hardest tasks. state of the art is not "it mostly works." plan your delegation budget around that.
the numbers
- Fable 5: 29.3% on FrontierCode Diamond
- Opus 4.8: 13.4%
- GPT-5.5: 5.7%
FrontierCode Diamond scores autonomous patches on real open-source repositories against held-out tests, not toy problems. 29.3% is the frontier, and 70.7% failure is what "frontier" means on the hardest slice of real work today.
what a realistic delegation budget looks like
if you hand off ten of your hardest problems, expect roughly seven to come back wrong, incomplete, or subtly broken. that is not a reason to stop delegating, it is a reason to delegate the right tier of task and verify the rest, per the four gates. the easier 80% of real work clears at a much higher rate than this number implies, FrontierCode Diamond is specifically the hardest slice.
one email when the next field note drops.
no course pitch, no daily emails. the notes, when they exist.
source: FrontierCode Diamond benchmark, Claude system card, page 256. last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork. related reading: can you trust what ai says it did?.