Claude Opus 4.8 Coding & Agent Benchmarks — SWE-bench Pro 69.2%, Weaknesses and All

The 30-second answer: Claude Opus 4.8 is at the front of the pack on coding, agentic, and real-world automation benchmarks. It scores 88.6% on SWE-bench Verified and 69.2% on the hardest tier, SWE-bench Pro — 10.6 points ahead of GPT-5.5. There is a trade-off, though: on single-task terminal work (Terminal-bench) and pure scientific reasoning (GPQA), … Read more

Claude Opus 4.8 Release — 5 Key Changes vs 4.7 (Launched May 28, 2026)

The 30-second answer: Claude Opus 4.8 is Anthropic’s new Opus model, released on May 28, 2026. The key changes versus 4.7 are a roughly one-quarter reduction in the rate of letting code defects through, an honesty improvement that avoids unsupported claims, and the addition of a fast mode that’s about 2.5x faster and 3x cheaper. … Read more

Claude Opus 4.8’s 1M Context — What a Solo Developer Actually Gains

The 30-second answer: Claude Opus 4.8 is Anthropic’s latest Opus-generation model, and it supports a 1M (one-million) token context window. For a solo developer, the biggest change is being able to load many files from a codebase and long documents at once and work without losing the thread. The official input price is $5 per … Read more