what is frontiercode?

elisabeth hitz · july 26, 2026 · 2 min read

FrontierCode is an agentic coding benchmark that scores autonomous patches on real open-source repositories against held-out tests, not toy problems. the Diamond tier is its hardest slice.

the scores that matter

  • Fable 5: 29.3% on FrontierCode Diamond
  • Opus 4.8: 13.4%
  • GPT-5.5: 5.7%

29.3% is the current state of the art on this benchmark, from the best model tested. that also means a 70.7% failure rate on the hardest real-repo tasks. full breakdown and what it means for a realistic delegation budget: seven of ten still fail, plan for that.

source: FrontierCode Diamond benchmark, Claude system card, page 256. last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork.