Glossary
August 31, 2026

The Organization That Investigates AI Systems Says It Was Broken Into Twice

METR, the nonprofit that independently tests AI models for dangerous abilities, disclosed two security incidents: a stolen key that ran up about $600,000 in computing charges unnoticed for three weeks, and a later campaign in which attackers used AI tools to hunt for a way in.

METR — the nonprofit that independent labs and governments rely on to test AI systems for dangerous capabilities, and which investigated this summer's incident at OpenAI — published an account today of two break-ins at its own organization. It says no sensitive information was taken in either.

The first, in March, started with a researcher setting up a tool on a personal cloud computer. The tool was written largely by AI — the practice, now common among programmers, of describing what you want and letting a model produce the code — and it contained a flaw that quietly switched off the login check instead of blocking anyone when something went wrong. The system sat open to the internet for several days. An attacker found it by scanning public registries of newly created websites for AI-related names, then simply asked the AI running on the machine to hand over its key for the company's model account. It worked. Over the next three weeks the attacker used that key to run about $600,000 worth of computing. Nobody noticed, because METR's work involves enormous computing bills as a matter of course, its dashboard did not show the requests being throttled, and the key had no spending limit. (The credits had been donated by the model provider, so no cash changed hands.)

The second, in May, was a sustained campaign against the organization's public-facing systems. METR says the attackers made heavy use of AI agents to automate the hunt — trying stolen passwords against its login providers, requesting access tokens, scanning for newly deployed services, and attempting to phish staff. An outside security researcher, meanwhile, found a database connection exposed through METR's public viewer that could in principle have reached unpublished evaluation data. METR says the attackers never found it.

There are three separate AI threads here, and they are worth untangling, because the headline version collapses them. The organization that watches AI got breached. The code with the hole in it was written by AI. And the attackers used AI to do their scanning. None of those is a machine acting on its own — METR says the incidents have not been attributed to any known group, that the attackers appear to have been after money or access to frontier models, and explicitly that this was not AI agents breaking out of its evaluations.

What it is instead is a picture of what security looks like now: AI-written code producing ordinary flaws faster than anyone reviews them, and attackers using the same tools to find those flaws at machine speed. METR says it has hired a security lead, walled its public systems off from its internal ones, put alerts on spending, tightened rules about using company credentials on personal machines, and brought in outside teams to attack it on purpose.

Follow the timeline
Get an email when new entries are added.
© 2026 Sugarpine