Glossary
December 20, 2024

OpenAI's Next Reasoning Program Scores Near Human Level on a Test Designed to Resist Memorization

On the last of its twelve days of announcements, OpenAI previewed o3, which solved 88 percent of puzzles on the ARC test — visual pattern problems that adults find easy and every earlier program failed — and a quarter of unpublished research-level math problems. It is not yet released.

OpenAI closed its twelve days of announcements today with a program it isn't releasing yet. It's called o3 — the company skipped o2 to avoid a trademark clash with a British phone carrier — and the results it showed are the reason people in the field spent the afternoon arguing.

The one that matters is a test called ARC. It consists of small visual puzzles: grids of colored squares where you have to work out the rule from a few examples and apply it to a new case. Most adults solve them easily. They were designed in 2019 specifically to be hard for AI, because they can't be memorized — each puzzle is new, and the training data doesn't help. For five years the best programs scored in the single digits or low teens. GPT-4o scores about 5 percent. o3 scored 76 percent using a normal amount of computing, and 88 percent when allowed to think at great length and expense. Human test-takers average about 85. The test's creator, François Chollet, called it "a genuine breakthrough" while noting that o3 still fails some puzzles a child would find trivial.

On a set of research-level math problems assembled by professional mathematicians and never published, where no previous program solved more than 2 percent, o3 solved 25 percent. On competitive coding it ranks among the top 200 human competitors in the world.

The caveats are real: the ARC result used enormous computing power — thousands of dollars per puzzle at the high setting — and OpenAI chose the tests. Release is planned for early next year, after safety testing. But the trend line, from o1 in September to this in December, is steep, and it is the trend line that people are arguing about.

Follow the timeline
Get an email when new entries are added.
© 2026 Sugarpine