rajveer (@rajveerbach) on X

X (formerly Twitter) ·

1 min read Original article ↗

@arrakis_ai

May 30

Claude Opus 4.8 has landed on DeepSWE Bench, posting a 58% Pass@1 and taking #2 overall behind GPT-5.5. It continues a broader trend: slightly behind on raw score, but among the most reliable and efficient coding models across recent benchmarks.