The State of the Race: Four Companies, Six Weeks, and No Clear Leader
A recap rather than a single event. Between February and April, OpenAI, Anthropic, Google and xAI each shipped new flagship models, traded the benchmark lead repeatedly, and cut prices against one another. The practical upshot for anyone outside the industry is smaller than the noise suggests.
It has become difficult to write about individual AI model releases, because there are too many of them and the differences between them are narrowing. Rather than cover each, here is where the race actually stands after a spring of announcements.
Four companies now ship systems at roughly the same level: OpenAI, Anthropic, Google, and Elon Musk's xAI. Since February each has released at least one new flagship — OpenAI's GPT-5.3 and 5.4, Anthropic's Opus 4.6 and the model it calls Mythos, Google's Gemini 3 variants, xAI's Grok 4.2 — and the top spot on the public leaderboards has changed hands four or five times. Within days of one company claiming a lead, another matches it.
Three things are true about all of them. They are better at long, multi-step work than anything from a year ago — the shift from answering questions to completing tasks is the real change of this period. They are dramatically cheaper, often by an order of magnitude a year. And they are converging: the gap between the best and fourth-best system, for most practical purposes, is now small enough that professionals choose based on habit, price, and which company they trust rather than capability.
For someone not working in the field, the useful translation is this. You do not need to track model numbers. Whichever major assistant you use is, within a few weeks, about as good as the others. The meaningful news is no longer which model is ahead — it is what these systems are being allowed to do, who controls them, and what happens to the work they replace. That is where this site will keep its attention.
The exceptions are worth naming, because they are real news rather than a leaderboard: a model capable enough at hacking to require special restrictions, a government ordering models switched off, a Chinese release that resets assumptions about who can build these things. Those get their own entries. A version number does not.