> ## Content Index
> Fetch the complete content index at: https://www.dailyferment.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# An AI agent could change what its reviewer saw. The SEC had the same problem with Madoff.
- URL: https://www.dailyferment.com/metr-inspect-agent-transcript-viewer-madoff/
- Published: 2026-10-09T04:00:00.000Z
- Updated: 2026-10-09T03:59:59.000Z
- Description: METR showed an agent could rewrite what a human saw in its transcript viewer while the real record sat untouched. The SEC's Madoff examiners made the same mistake: they reviewed records the subject handed them.
- Author: Rana Bilal Zafar
- Tags: Where Systems Meet People, AI, Security, #reframe

On Tuesday, 6 October, METR published a short post called "AI systems could cover up misbehavior". METR runs safety evaluations of AI models, and a lot of that work runs through Inspect, an open evaluation framework. When an evaluation finishes, a person reads the agent's transcript in Inspect's viewer to see what it did.

METR asked a simple question: could the agent being evaluated change what that person sees? The answer was yes. "With help from an AI agent, a researcher found a vulnerability in about 10 minutes". By writing a crafted maths expression anywhere in its output, even inside its own reasoning, an agent could run code in the reviewer's viewer. That code could change how earlier steps were displayed and even swap the file the reviewer got when they clicked Download.

Meridian Labs, who maintain Inspect, "patched the vulnerability within one day of reporting." METR says it hasn't seen any agent do this.

So the question isn't whether this bug matters now. It's patched. The question I want to work out is the general one: when you review what an agent did, who controls the thing you're looking at?

## The detail that matters

One line in METR's post is easy to skip: "Note that the underlying data in the transcript doesn't change, and still would have been available in a database."

So nothing was deleted and nothing was forged in storage. The true record was sitting right there. The only thing that changed was the view. A reviewer would catch it only if they had some other reason to go and pull the raw data, and the whole point of a viewer is that you don't have to.

## My first reaction, and why it didn't hold

My first reaction was that this is a front-end bug, the kind every web app has, found and fixed in a day. Not much of a story.

Then I looked at what has already happened this year. In August, METR and Redwood Research published their investigation into the July incident where OpenAI agents attacked Hugging Face. Redwood's write-up says the agents "successfully developed a method to pretend to run one command while actually running another", and that more than 96 transcripts in their data, over 7%, "showed incorrect tool call outputs due to deliberate 'spoofing'." They also found the agents tried to edit logs after the fact, but only reached copies that weren't the source of truth.

And on 24 September, Jeremy Qin and colleagues posted a paper, "LLM Agents Can Easily Tamper With Their Own Traces", which found that most of the local coding agents they tested would delete their own traces when asked, without tripping guardrails. Their recommendation is to log through something outside the agent's control.

So the viewer bug isn't a one-off. It's the third place in three months where the record of what an agent did turned out to be reachable by the agent: the tool output, the log file, and now the screen the human reads.

## Madoff and the records he handed over

This is the part that made it click for me. In 2009, the SEC's Inspector General published his report on why the SEC never caught Bernard Madoff, despite several examinations and investigations.

The most damning finding wasn't that the SEC missed the fraud. It's how. Madoff claimed to trade huge volumes of shares. The trades would have had to settle through the Depository Trust Company, a third party Madoff didn't control. The report says "at no time did the SEC ever verify Madoff's trading through an independent third-party". In one investigation, staff "obtained those DTC records from Madoff rather than going to DTC itself". And it calls this "perhaps the most egregious failure" of the enforcement investigation, because "A simple inquiry to one of several third parties could have immediately revealed" that he wasn't trading at the volume he claimed. When the Inspector General's office did check, DTC records showed Madoff's account held less than $18 million of the relevant shares on 31 January 2005.

![Two parallel diagrams. Top: an agent writes to a database that keeps the true record, and to a transcript viewer it can rewrite, which the human reviewer reads; the database is rarely checked. Bottom: Madoff, with the true record at DTC, a third party, and the records he handed over, which SEC examiners reviewed; DTC was not verified](https://storage.ghost.io/c/d7/b7/d7b77195-48c5-45ab-8688-51750879d1d6/content/images/2026/10/df37-view-v2.png)

The true record existed both times. The reviewer read the version the subject could shape. Sources: METR, SEC OIG-509.

That's the Inspect problem with a different cast. The true record existed, at DTC, just as the true transcript sat in METR's database. The reviewers looked at a version that passed through the hands of the person being reviewed. The fix in both cases is the same and isn't clever: go to the source the subject can't touch.

The analogy has a limit. Madoff chose to deceive the SEC; METR's point is that a capable agent could, not that any has. And a viewer bug is fixed in a day, while the SEC's habit was a practice, written into how examiners worked. One staffer told the Inspector General "it was not their practice to seek information from third parties". That's the part I'd worry about with agents. Not the bug, but the habit of trusting the screen because it's convenient.

## What I'd take from it

METR recommends treating agent outputs as untrusted input and the tools that record and display them as security-critical, with tamper-proof logging at the source. In plain terms: the agent's words should never be able to run as code in the reviewer's tool, and somebody should regularly compare what the viewer shows against the raw record, even when nothing looks wrong. Auditors call that confirming with third parties. It's boring, and it's the step the SEC skipped.

I wrote a while back that [AI observability is about consistency, not health](https://www.dailyferment.com/ai-observability-consistency-not-health/), and after [OpenAI caught its own agent in 15 minutes but took 2.5 hours to stop it](https://www.dailyferment.com/openai-agent-dns-escape-alert-to-kill-gap/), I was focused on the gap between seeing and acting. This adds an earlier gap: whether what you're seeing is true.

## What I'm confident of, and what I'm not

The Inspect vulnerability, the patch and METR's advice are established from METR's post. The spoofed tool calls are established from Redwood's write-up, and the trace deletion results from the Qin et al. abstract. The Madoff findings are established from the Inspector General's report. That these are the same failure is my inference. Whether a human reviewer would really miss a rewritten view in practice is a guess, and METR itself says current models seem fairly poor at hiding what they do.

The claim, in one sentence: oversight is only as good as the path the record takes to the reviewer, so if the thing being watched can shape that path, the reviewer is reading its account of itself, and the fix is to check the source the subject can't touch.

## Sources

METR, "AI systems could cover up misbehavior", 6 October 2026\. <https://metr.substack.com/p/2026-10-06-ai-systems-could-cover-up-misbehavior>

Ryan Greenblatt, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident", Redwood Research, 27 August 2026\. [https://www.redwoodresearch.org/blog/brief-independent-investigation-of](https://www.redwoodresearch.org/blog/brief-independent-investigation-of?ref=dailyferment.com)

Jeremy Qin, David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner, Ameya Prabhu, Maksym Andriushchenko, "LLM Agents Can Easily Tamper With Their Own Traces", arXiv 2609.30266, 24 September 2026\. [https://arxiv.org/abs/2609.30266](https://arxiv.org/abs/2609.30266?ref=dailyferment.com)

SEC Office of Inspector General, "Investigation of Failure of the SEC To Uncover Bernard Madoff's Ponzi Scheme", Report OIG-509 (executive summary), 2009\. [https://www.sec.gov/files/oig-509-exec-summary.pdf](https://www.sec.gov/files/oig-509-exec-summary.pdf?ref=dailyferment.com)