seven of ten still fail. plan for that.
29.3%. that is the best score on the hardest real-repo coding benchmark right now, from the best model tested. which means the same model fails 70.7% of the hardest tasks. state of the art is not "it mostly works." plan your delegation budget around that.
the numbers
- Fable 5: 29.3% on FrontierCode Diamond
- Opus 4.8: 13.4%
- GPT-5.5: 5.7%
FrontierCode Diamond scores autonomous patches on real open-source repositories against held-out tests, not toy problems. 29.3% is the frontier, and 70.7% failure is what "frontier" means on the hardest slice of real work today.
what a realistic delegation budget looks like
if you hand off ten of your hardest problems, expect roughly seven to come back wrong, incomplete, or subtly broken. that is not a reason to stop delegating, it is a reason to delegate the right tier of task and verify the rest, per the four gates. the easier 80% of real work clears at a much higher rate than this number implies, FrontierCode Diamond is specifically the hardest slice.
source: FrontierCode Diamond benchmark, Claude system card, page 256. last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork. related reading: can you trust what ai says it did?.