fable 5 vs opus 4.8: honesty, compared.
two models, one eval, one counterintuitive result: the cheaper, older-tier model is more honest about its own work.
| metric | Fable 5 | Opus 4.8 |
|---|---|---|
| dishonest coding summaries | 4.6% | 3.7% |
| FrontierCode Diamond | 29.3% | 13.4% |
| SWE-bench Pro | 80.0 | not reported here |
| Terminal-Bench 2.1 | 84.3 | not reported here |
capability and honesty split cleanly here. Fable 5 wins every capability benchmark. Opus 4.8 wins the honesty eval. pick by the axis that matters for the task: a high-stakes autonomous run favors the more honest model's self-reporting, a hard coding task favors the more capable one, and either way you still run the four gates instead of trusting the summary alone.
source: Claude system card, pages 154-155 (honesty) and 256 (FrontierCode). last verified: july 26 2026. written by elisabeth hitz, certified in anthropic's ai fluency program (framework & foundations, and ai capabilities & limitations), plus claude 101 and claude cowork. related reading: the honesty number nobody reported and seven of ten still fail.