Claude Opus 4.8 Coding & Agent Benchmarks — SWE-bench Pro 69.2%, Weaknesses and All
The 30-second answer: Claude Opus 4.8 is at the front of the pack on coding, agentic, and real-world automation benchmarks. It scores 88.6% on SWE-bench Verified and 69.2% on the hardest tier, SWE-bench Pro — 10.6 points ahead of GPT-5.5. There is a trade-off, though: on single-task terminal work (Terminal-bench) and pure scientific reasoning (GPQA), … Read more