Member-only story
AI Benchmarks Are a Joke. Grok 4’s Big Flop Is The Proof.
From sky-high leaderboard scores to “WTF is this?” results, a new wave of user backlash reveals that AI may be acing all the wrong exams. Is it time we rethink how we grade intelligence itself?
In the past few years, the AI world has felt like a nonstop race to the top. Every few months, a new leader takes the spotlight. Companies like OpenAI, Google, Meta, and now xAI are locked in a constant competition.. trying to beat each other on benchmark scores, leaderboards, and test results. It’s the tech equivalent of the Olympics, and we’re all supposed to be on the edge of our seats.
The latest contender to climb the mountain was Grok 4, Elon Musk’s brainchild, which was hyped up as the “smartest AI in the world.” And if you just looked at the scorecards, you’d believe it. It crushed benchmarks like GPQA, Humanity’s Last Exam, and a bunch of other tests with names that sound like they were cooked up in a sci-fi nerd’s room.
But then, a funny thing happened. The AI escaped the lab and landed in the hands of real people. You know, the ones who actually need to get stuff done.
