Anthropic's Mid-Sized Chatbot Now Beats Its Own Flagship, and Its Rivals', at Twice the Speed
Claude 3.5 Sonnet outscores GPT-4o and Google's Gemini on most tests and works alongside you in a side panel that shows code, documents and diagrams as it writes them. It's free.
Anthropic released Claude 3.5 Sonnet today, and the naming undersells it. Sonnet is the company's middle size, but this version beats Claude 3 Opus — the large, expensive model it released three months ago — on every benchmark Anthropic publishes, at twice the speed and a fifth the cost. It also edges out OpenAI's GPT-4o and Google's best on most of them. It is free to use on the company's website.
The feature people are talking about is called Artifacts. Ask Claude to write a program, draft a document, or make a chart, and it appears in a panel beside the conversation, where you can watch it take shape, edit it, and run it. The chatbot stops being a window you type into and becomes something closer to a workspace. Anthropic says this is where it is headed: toward programs that do work alongside people rather than just answering them.
The pace is the other story. Three months ago Anthropic's best model was the best in the world for about ten weeks; then OpenAI matched it; now Anthropic has passed both with a smaller, cheaper program. The three leading labs are trading the lead every few months, and the improvements are arriving as much from better training as from bigger machines.
Independent testers have generally confirmed the company's claims about the last two releases, so the benchmarks are likely to hold. What they can't measure is the thing Anthropic is quietly best at: many programmers now say it is simply the one they prefer to work with.