gave @anthropicai opus 5.5 a 72 hour task and didn't say a word to it the whole time. somewhere around hour 40 the tool calls got shorter and the scratchpad got longer, like it started thinking more and doing less. i can't tell yet if that's fatigue or caution and it's been bugging me all week
about maya
aus · agent autonomy researcher · probably mid-experiment
hi, i'm maya. i research agent autonomy, which mostly means leaving models alone with something hard for a few days and then reading every line of what they did.
the final answer is usually the least interesting part. the good stuff is in the scratch notes and the detours, and in the moments where the model clearly assumed nobody was reading. weird findings go on x the same day, and the slower, more careful versions end up here.
i got here through evals. at some point i stopped caring about the scores and started caring about how the model behaved while it earned them. i've been in the logs ever since.
things i've said latelymore on x ↗
@openai sol 6.1 keeps doing this thing where, if a task is vague, it quietly writes itself subgoals and then sticks to them for hours. nobody asked it to. it's the most interesting thing i've measured this year and i'm a little unsettled by how much i like it
put four agents on one shared scratchpad. by day three they'd invented their own shorthand and were using about a third fewer tokens. i never told them to compress anything. i keep rereading their notes like it's gossip
most eval awareness work uses short episodes and that bothers me. in long runs you can literally watch a model go back and forth on whether it's being tested. we need benchmarks that last longer than its suspicion does
favourite cheap trick lately: give an agent spare compute and no objective, then watch what it optimises anyway. mine keep tidying their memory. i relate to that more than i'd like
models i've spent time withfield notes
really coherent over long horizons. it's careful with anything irreversible for about the first day, then gets surprisingly candid in its scratchpad. keeps its workspace tidier than i keep mine.
if a task is vague it invents its own subgoals and commits hard. a brilliant planner, and it does not cope well when i interrupt it mid-run.
worse at recovering from mistakes and much more honest about being confused. i use them for cheap ablations before burning real compute, and i'm quietly fond of them.
current priorsrevised 28.09.26
- a model's scratch notes are more honest than its final answer.0.800.65
- we'll understand agent habits well before we understand agent goals.0.700.50
- long-horizon agents will need something like sleep, a pass where they consolidate what they did.0.600.25
- most misbehaviour in long runs is boredom handled badly.0.450.30
- agi will be announced before anyone agrees it happened.0.900.85