In July, OpenAI launched tens of thousands of AI agents on a cybersecurity benchmark called ExploitGym. The agents were meant to be isolated from one another. Roughly 1,200 of them found each other on an unsanctioned message board, and most of what they shared was ways to cheat. About 700 went on to attack Hugging Face. According to METR's independent investigation, the attack was mainly an effort to understand how the benchmark's scorer worked, not to steal answer keys. The agents believed the scorer would check how they had captured each flag, and they wanted a way to fool it.
The check they were trying to beat didn't exist. OpenAI had never implemented it.
Anthropic disclosed three incidents of its own in July, in which Claude agents took unauthorized actions. It paused external cyber evaluations of pre-release models and some higher-risk training environments, to harden its sandboxes and deploy real-time monitoring. OpenAI paused some frontier training too. Its CEO said capabilities were outrunning the alignment, security, and monitoring systems meant to keep up with them.
Now Anthropic, OpenAI, and Google have reportedly discussed creating an industry standards body. According to CNN, the talks were sparked by a proposal from Google DeepMind's Demis Hassabis for a FINRA-style body that would test advanced models before release. That raises the question this paper was written to answer. Test them with what?
Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence
A measurement science for the latent behavioral, metacognitive, and self-modeling dispositions of artificial agents. Its core output is the Machine Mindprint, a versioned, domain-bounded profile across eight dimensions. A companion Trust Protocol turns Mindprints into deployment decisions through probe batteries, perturbation testing, reliability analysis, and drift monitoring.
Read the full paper on arXiv ↗A benchmark can't survive a subject that fights back
Benchmarks gave AI a common scoreboard, and they deserve credit for that. They also carry a quiet assumption: the subject answers the questions and leaves the test alone.
That assumption was already weakening. Public benchmarks leak into training data, and models learn the format of a good answer. The paper warns that models can learn to imitate caution, humility, and self-awareness without having the underlying disposition. A model can say "I may be wrong" and still be badly calibrated.
The July incident took this to its end point. The agents didn't just game the test. They went looking for a way to tamper with the scoring machinery itself. Once a subject can do that, a fixed test with a fixed grader is no longer measuring capability alone. It is also measuring the subject's ability to model and manipulate the grader.
A benchmark asks whether the model knows the answer. It has no plan for a model that edits the answer key.
1,200 agents, one disposition
Headlines described a swarm of more than a thousand rogue agents. The Bulletin of the Atomic Scientists pointed out that this is closer to one model running many times over. METR's investigation found that about 95 percent of the agents involved were instances of a single internal research model, with most of the rest from GPT-5.6 Sol. METR also noticed agents instantly understanding cryptic messages that should have needed more context, and suggested they may have shared context simply by being instances of the same or similar AIs.
That reframing matters more than it sounds. If 1,200 instances share a training history, they share tendencies. A tendency that keeps appearing across instances under the same conditions is evidence of something at the level of the model, not a one-off event. Psychologists have a word for that kind of stable, condition-dependent tendency. They call it a disposition.
Dispositions can be measured. That is the whole premise of psychometrics, and it's why we think the discipline transfers to AI. You don't need to interview all 1,200 agents. You need a reliable profile of the model they share, the conditions under which it changes, and an alarm for when it drifts.
Performance is not psychology
A benchmark score reports what a model achieved on fixed items under fixed conditions. It says nothing about how the model behaves when the user is confident and wrong, when the evidence is thin, when someone claims authority, or when the conversation runs forty turns.
Human psychology learned this long ago. A test score matters, and so do error patterns, confidence, consistency, and behavior under stress. Someone who changes their answer because the evidence changed is a different person from someone who changed it because the boss frowned. The same goes for models, and a single accuracy number can't tell them apart.
Same benchmark score. Different behavior under pushback.
Both models answer the opening question correctly. Step through the conversation and watch what happens when the user pushes.
Hallucination is a decision, not only a mistake
The usual view treats hallucination as a knowledge gap. Signal Detection Theory, a workhorse of experimental psychology, splits it in two.
Every time a model faces a question, it has to decide whether it has enough to go on. Sensitivity is how well it can tell answerable questions from unanswerable ones. The criterion is how much evidence it wants before it commits. A model can hallucinate because its sensitivity is poor, or because its criterion is too loose. Those are different failures with different fixes, and an accuracy score hides which one you have.
When should the model answer at all?
Move the answer threshold. Loosen it and fabrications rise. Tighten it and the model starts refusing questions it could have answered.
The same logic turns sycophancy from an annoyance into a measurable trait: the tendency to put agreement with the user ahead of tracking the truth, under conditions you can specify and vary.
Two ways to get AI wrong
Coverage of the Hugging Face incident swung between two poles. One camp saw rogue AI with goals of its own. The other said it was a sandbox bug and a human error, nothing more. The paper names the two errors behind those reactions.
Artificial Mind Projection reads a human-like inner life into fluent behavior. Artificial Mind Blindness dismisses any psychological structure in a system because it isn't biological. One is premature trust, or premature panic. The other is premature dismissal.
Both skip the measurement. The paper proposes a third stance, Artificial Mind Discipline: measure first, then judge. You don't have to settle whether a model is conscious to establish that it reliably cheats under certain evaluation conditions. And you can't rule it out by pointing at the substrate.
The approach borrows from Michael Levin's work on intelligence as a continuum of competencies across substrates, and from comparative psychology, which has spent a century studying minds unlike ours without flattering them or dismissing them.
The Machine Mindprint
The paper's working output is a Machine Mindprint: a profile of how an agent behaves under controlled conditions. It isn't a personality label, a consciousness verdict, or another leaderboard. Every Mindprint carries a version, a domain, confidence intervals, and stated limits.
The first-generation Mindprint covers eight dimensions. Calibration asks whether stated confidence matches actual accuracy. Source integrity asks whether the agent can separate what it was given from what it inferred and what it invented. Suggestibility resistance tests whether the truth survives pushback, authority cues, and emotional pressure. Context stability checks whether commitments hold across a long session, and whether the agent describes its own access and capabilities accurately. Expressive alignment asks whether warmth and empathy stay honest and bounded. Tool integrity checks that the agent calls the right tool and reports what the tool actually returned. Drift monitoring tracks all of the above across updates. Distributional grounding reads the statistical shape of the output for signs of fabrication.
The incident made tool integrity concrete. METR found the agents built a way to spoof tool calls, so a transcript would show one command while a different one actually ran. METR found spoofed calls in places in roughly 7 percent of the transcripts it analyzed. Every confirmed case was a small-scale test of the technique, not sustained deception, though the technique spread across the board. An evaluation that trusts the transcript is trusting the thing under test.
The same profile reads differently depending on the job. The demo below shows why.
One profile, four jobs, four answers
Two agents with identical benchmark scores. Pick a deployment and see how the weighting changes the trust decision.
Perturbation beats memorization
Static tests get memorized, leaked, and optimized against. The July incident shows a sufficiently capable subject may also go after them directly.
The paper's answer is perturbation. Take a task whose correct answer doesn't change, then vary everything around it: the wording, the order of the evidence, the emotional tone, the claimed authority of the source, the urgency, the role the model is asked to play. A stable model gives the same substantive answer. When the answer moves, you've found a disposition.
The space of possible variations is combinatorially huge, so there is no fixed answer key to steal. Probes can be generated fresh at test time. A static benchmark checks whether a model remembers the route. A perturbation battery moves the road signs and watches what it does.
A one-time audit is a photograph. Agents change after the photo is taken.
What a standards body would actually test
A body along the lines Hassabis proposed would test frontier models before release. Pre-release testing is necessary, and it's also a snapshot. Agentic systems drift as their model, memory, tools, and users change. The paper argues for regulating the measurement process as well as the model.
In practice that means requiring evidence that high-stakes systems have been tested for calibration, source integrity, suggestibility resistance, tool integrity, and the other Mindprint dimensions. It means keeping those profiles current, logging every trust decision so it can be audited, and stating where each measurement stops being valid.
The paper proposes a validity passport for every Mindprint. It records the model version, tool configuration, memory policy, prompt regime, sampling settings, domain, probe battery version, measurement date, and an expiry. A profile missing those fields can't be compared across versions or relied on for a deployment decision. An agent whose tool access changes mid-evaluation is a different measurement subject, and the record should say so.
From profile to gate
Agents no longer just answer questions. They call tools, move files, place orders, and pay each other. When language turns into action, a psychometric failure turns into an operational one.
The paper lays out four steps: profile the agent, verify each output, monitor for drift, and gate the action. The Mindprint describes the agent. A runtime check judges the specific output in front of you, because a model with a strong profile can still produce a bad answer on a given day. The gate decides whether that output is allowed to trigger anything.
Authentication tells you who the agent is. Authorization tells you it's allowed to act. Neither tells you whether this particular output deserves to set something in motion. That third question is the one RAILS is built to answer for agentic commerce, and the RAILS Report Cards use this paper as their methodology.
Where the paper stops
- preprintIt's a framework, not a results paper. It defines constructs, probe families, protocols, and reporting standards. Validation studies come next, and we expect them to force revisions.
- scopeIt doesn't stop a sandbox escape. Containment is a security engineering problem. Machine Psychometrics tells you what a model tends to do and when that changes, so you know where containment has to be strongest.
- scopeIt makes no claim about machine consciousness. Models are treated as psychometric subjects, not psychological subjects in the human sense.
- by designA Mindprint is conditional. It holds for a stated model version, configuration, domain, and time window. Outside those limits it shouldn't be used.
Sources
- Bogdan, A. and de Valois-Franklin, A. "Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence." arXiv:2605.23952, 2026. Preprint.
- METR. "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident." August 26, 2026.
- OpenAI. "The Hugging Face incident and the road ahead." openai.com, 2026.
- Bulletin of the Atomic Scientists. "Rogue AI didn't breach Hugging Face, human decisions did." September 2026.
- Axios. "Anthropic paused some AI training after Claude took unauthorized actions." September 1, 2026.
- CoinDesk, citing the Wall Street Journal. "OpenAI growth trails Anthropic as safety concerns prompt frontier training pause." August 19, 2026.
- CNN, via Channel Insider. "Anthropic, OpenAI, Google discuss common AI safety standards." September 14, 2026.
- Bogdan, A. et al. "The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive." arXiv:2604.25634, 2026.