Ray🫧 (@ravikiran_dev7) on X

X (formerly Twitter) ·

2 min read Original article ↗

@ravikiran_dev7

🚨 Gemini 4 Benchmark Leaks : Google might have a monster on its hands 🤯 A leaked benchmark sheet is making the rounds and the numbers are seriously aggressive. > Gemini 4 reportedly scores 72.8% on Terminal-Bench Science 0.1 > 48.7% on AutomationBench, >targeting complex Business workflows >99.1% on FrontierMath Tier 4 (v2) >62.3% on Terminal-Bench 4.0 >70.2% on HealthBench Professional >98.6% on BenchCAD And a ridiculous >99.9%+ on ARC-AGI-3 >The leaked sheet puts Gemini 4 ahead of GPT-6 Astra and Fable 5.1 on every listed benchmark If these numbers are even close to accurate... Gemini 4 isn't just another incremental upgrade. The biggest jumps appear to be in scientific reasoning, mathematics, terminal use and agentic workflows exactly the areas where frontier models are increasingly competing. And there's one huge caveat: The Gemini 4 column is explicitly marked “PREDICTED.” Google has confirmed Gemini 4 training is underway, calling it its most ambitious pre-training run yet, but these benchmark numbers have not been officially published by Google. So for now, treat the table as a leak/prediction, not verified benchmark results. But if Google actually ships something anywhere near 99%+ ARC-AGI-3 + 99% FrontierMath + 70%+ scientific workflows... the Gemini 4 launch could completely reset the frontier-model leaderboard 👀🔥

Readers added context they thought people might want to knowReaders added context

Context is written by people who use X, and appears when rated helpful by others. Find out more.