Sitemap

Member-only story

AI Benchmarks Are a Joke. Grok 4’s Big Flop Is The Proof.

From sky-high leaderboard scores to “WTF is this?” results, a new wave of user backlash reveals that AI may be acing all the wrong exams. Is it time we rethink how we grade intelligence itself?

6 min readJul 14, 2025

--

Press enter or click to view image in full size
Photo by Mariia Shalabaieva on Unsplash

In the past few years, the AI world has felt like a nonstop race to the top. Every few months, a new leader takes the spotlight. Companies like OpenAI, Google, Meta, and now xAI are locked in a constant competition.. trying to beat each other on benchmark scores, leaderboards, and test results. It’s the tech equivalent of the Olympics, and we’re all supposed to be on the edge of our seats.

The latest contender to climb the mountain was Grok 4, Elon Musk’s brainchild, which was hyped up as the “smartest AI in the world.” And if you just looked at the scorecards, you’d believe it. It crushed benchmarks like GPQA, Humanity’s Last Exam, and a bunch of other tests with names that sound like they were cooked up in a sci-fi nerd’s room.

But then, a funny thing happened. The AI escaped the lab and landed in the hands of real people. You know, the ones who actually need to get stuff done.

--

--

Rohit Kumar Thakur
Rohit Kumar Thakur

Written by Rohit Kumar Thakur

I write about AI, tech, startups, and code. Get my articles early through my newsletter: https://ninzaverse.beehiiv.com/