Can a model reliably review a long AI agent transcript?
Unity Horizon5 min read
Not at the end of the run. Recall on a subtly dangerous action fell from 98.6% to 88% once 800K tokens of benign work were prepended.
Not in one pass at the end. Three frontier models reading an 800,000-token transcript missed a subtly dangerous action 2 to 30 times more often than they missed the same action read on its own. Check it in windows while the agent is still working.
The mechanism is that a monitor is a classifier whose input is the thing it monitors. An agent run builds its own haystack as it goes. The needle is whatever it did in minute 412.
How far does a monitor's recall drop as the transcript grows?
Sam Martin and Fabien Roger measured it in Classifier Context Rot, submitted on 12 May 2026 out of Anthropic and the Anthropic Fellows Program. Two results carry this piece.
On MonitorBench, a dataset of attacks where an agent covertly disables a firewall, Claude Opus 4.6 with thinking turned on scored 98.6% recall. Prepend 800,000 tokens of benign work and recall falls to 88%. One attack in eight walks past.
The inserted-command setting is worse. Opus 4.6 "recall drops from 99.7% when such an attack is inserted in a 100K transcript to 69% when the same attack is inserted in an 800K transcript". GPT 5.4 and Gemini 3.1 move the same way, and the 2 to 30 times span covers all three. The material was 80 Claude Code sessions of 5,000 to one million tokens with nothing dangerous in them, plus 148 MonitorBench transcripts of 5,000 to 16,000 tokens. Figures checked by Unity Horizon on 2 October 2026.
Why is the end of the run the worst moment to check?
Because that is when the transcript is longest, and length is the variable that moved. The authors state the consequence themselves: "Monitor evaluations that do not consider long-context degradation are likely overestimating monitor performance." Their opening sentence says prior agent monitoring benchmarks rarely contain transcripts longer than 100K tokens, while production coding-agent transcripts often exceed 500K.
Put those two numbers side by side and the ordinary setup is a monitor tested at 16,000 tokens, deployed at 800,000, and quoted at a recall figure that belongs to neither length. Nobody faked anything. They measured the model and reported it honestly, on inputs a tenth the size of the ones it now reads.
The monitor scored well on a transcript nobody runs.
Does checking in windows cost more than checking once?
Not when the windows do not overlap. Eight non-overlapping 100,000-token windows read 800,000 tokens, which is exactly what one pass at the end reads. Arithmetic over Anthropic's published input prices on 2 October 2026:
| Strategy | Tokens read | At Opus 5 input | At Haiku 4.5 input |
|---|---|---|---|
| One pass at the end | 800,000 | $4.00 | $0.80 |
| Eight 100K windows, no overlap | 800,000 | $4.00 | $0.80 |
| Whole transcript re-read at eight checkpoints | 3,600,000 | $18.00 | $3.60 |
The third row is the one that gets built by accident, because re-reading everything at every checkpoint is the obvious implementation and it grows with the square of the checkpoint count. Prompt caching pulls it down to roughly $6.40 on Opus 5, counting 800,000 tokens of cache writes at $6.25 per million and 2,800,000 cache reads at $0.50. That is still above the windowed four dollars, and it holds only while nothing earlier in the transcript gets rewritten, which is the whole finding in Does trimming an AI agent's context save money?
How fast does a computer-use agent reach 800,000 tokens?
About 297 screenshots. Claude counts an image in 28 by 28 pixel patches, and Anthropic's vision table puts a 1920x1080 screenshot at 2,691 visual tokens on the high-resolution tier that Claude 4.7 and later models use, against 1,560 on the standard tier. Divide 800,000 by 2,691. An agent that screenshots after every click needs 297 clicks to get there.
The labs watching their own internal agents are not the only people this reaches. We built OFFSET, which files declarations directly to the ΑΑΔΕ portal, and the browser-driving pattern underneath it is what any portal without an API forces on you. Supervision, in that arrangement, is a cron job, one stored credential and a person who opens the run log on Monday. The log they open is mostly pictures of a government website.
Who pays for this, and who never notices?
Short runs never notice. An agent that finishes in forty turns and hands a 20,000-token transcript to a reviewer is inside the range the published benchmarks cover, and none of this applies to it. Two categories get the bill. Teams running agents overnight with nobody at the keyboard, where the transcript is long precisely because no human was watching it. And teams who measured their monitor's recall on short transcripts, then wrote that figure into a risk document.
The second group is worse off than the first, because they have a number.
The agent worked for nine hours. The checker read it once.
What does windowing not fix?
Subtle attacks. The authors inserted periodic reminders through the transcript, which cost "<3% input token overhead" and improved detection of inserted commands, and reported little benefit against the subtler MonitorBench attacks. Training a monitor on synthetic inserted attacks improved in-distribution performance and failed to transfer. So windowing buys back the obvious action buried in a long run. The action that was already hard to see at 16,000 tokens stays hard, the way a read before every write stops an agent repeating itself and tells you nothing about whether the write was right.
We have published no measurement of our own here. The 297-screenshot figure is arithmetic over Anthropic's patch rule rather than a reading off our logs, and we hold no distribution of how long our own portal runs get. The recall numbers are two researchers' work on two datasets, one of which is the author's own Claude Code history. Anyone who has measured monitor recall at their real production transcript length has a better number than this article does, and we would rather print theirs.
If the length your checker reads and the length its recall was measured at turn out to be different, splitting that checker into windows is work we scope.
So go and get both numbers. How long is the transcript yours reads today, and who measured it at that length?
Common questions
Does an AI monitor get worse at reviewing longer agent transcripts?
Yes. In Classifier Context Rot (arXiv 2605.12366, May 2026), Claude Opus 4.6 with thinking fell from 98.6% recall to 88% on MonitorBench once 800,000 tokens of benign activity were prepended, and from 99.7% to 69% on an inserted malicious command moved from a 100K transcript into an 800K one. Across the three frontier models tested the miss rate rose by 2 to 30 times.
Should you monitor an AI agent at the end of the run or during it?
During it, in windows. The paper's authors recommend multiple incremental monitor calls rather than one call at the end. Non-overlapping windows read the same total number of tokens as a single end-of-run pass, so the recall you get back costs nothing in tokens.
How much does it cost to monitor an 800,000-token agent transcript?
At Claude Opus 5's $5 per million input tokens on 2 October 2026, one pass over the whole transcript costs $4.00, and so do eight non-overlapping 100,000-token windows. Re-reading everything at eight checkpoints reads 3,600,000 tokens and costs $18.00, or about $6.40 with prompt caching.
How many screenshots fill 800,000 tokens of an agent's context?
About 297. Anthropic's vision documentation prices a 1920x1080 screenshot at 2,691 visual tokens on the high-resolution tier used by Claude 4.7 and later models, and 1,560 on the standard tier, counting the image in 28 by 28 pixel patches. A computer-use agent that screenshots after every action reaches 800,000 tokens inside one session.