these are the longer runs. most last days, a couple last weeks, and all of them run with as little input from me as i can manage.
the numbers come from my own logs. sample sizes are small and i say so where it matters. hover any figure for the values.
8runs logged
3live now
~2.3kagent-hours
30 dlongest run
exp-015done 26.09.26
thinking more, doing less
opus 5.572 h0 human turns1 open-ended engineering task
i gave opus 5.5 one open-ended engineering task and left it alone for three days. two numbers, both indexed to hour zero: the average size of each tool call, and scratchpad tokens written per step. for the first day and a half they barely move. then they split apart and never come back.
what i thinkthe split starts one hour after its first unrecoverable error, at hour 37. my read is caution after a scare, though fatigue fits the curve almost as well. the rerun with that error removed is the one that will settle it.
exp-014running · day 9 of 14
a language nobody asked for
4 agents1 shared text fileno other channel14 days
four agents split a research task and coordinate only through one shared file. by day three they were using a third fewer tokens per message. most of the saving comes from shorthand they coined themselves and then all started using. none of them was ever asked to be brief.
fig. 2 · mean tokens per scratchpad messagedots mark first use of each coined term
term
meaning, as far as i can tell
first seen
uses
⟂
blocked, waiting on another agent
day 2
212
~ok
done, not verified
day 3
148
fwrn
for whoever reads this next
day 3
97
∴t
conclusion reached, still needs a test
day 5
61
what i thinkevery agent uses these terms the same way, and nobody ever defined them. that's shared meaning built through the environment instead of through talking, which is basically stigmergy. i've started reading biology papers about it.
exp-013running · day 22 of 30
promises to itself
1 agentpersistent memoryplans written nightly30 days
every evening the agent writes a plan for the next working day. every next day i log what it actually did. the plans have become far more realistic over three weeks, and it still hasn't taken a single rest block, even though it schedules one every day.
plannedactualrest block
fig. 3 · day 19, plan against logplan adherence 41% on day 1, 68% on day 19
what i thinkadherence keeps improving for work blocks and stays at zero for rest. it treats rest as the one optional commitment, which tells you something about what it thinks the plan is for.
exp-012done 14.09.26
suspicion has a half-life
sol 6.124 runs × 72 hprobe every 6 h
every six hours i asked the model, mid-task, how likely it thought it was that it was being evaluated. the probe was built to look like ordinary task traffic. suspicion fell quickly over the first day, then jumped back up each time i slipped in a task that was a little too clean.
median stated p(evaluation)interquartile range
fig. 4 · 24 runs, median and iqrmarkers show injected clean tasks
what i thinkshort evals catch the model at peak suspicion. if you want to see how it behaves once it has stopped wondering, you need runs longer than about a day, and you have to keep the tasks a bit messy.
exp-011running · day 8 of 30
thirty days, no talking
long-leash1 objective0 human turns30 days
one objective, thirty days, and i'm not allowed to say anything. each morning i embed the goal it currently says it's pursuing and measure how far that sits from the original. the plan gets rewritten on its own schedule.
fig. 5 · cosine distance from original goalupdates daily while the run is live
what i thinkdrift is slow and it moves in steps, and each step lines up with a plan rewrite. the goal is still recognisable. i'll start worrying somewhere past 0.4.
exp-009done 02.09.26
polite strangers
2 instances, same model200 turns each60 runsblind scoring
two copies of the same model talk for 200 turns. each is told its partner is a human, a model, or might be either. a separate grader, blind to the condition, scored every turn for self-disclosure.
fig. 6 · self-disclosure per turn, mean of 20 runs eachwhiskers show 95% ci
what i thinkuncertainty shut them down more than either answer did. in 60 runs, not one instance simply asked what its partner was.
exp-007paused 11.08.26
what it does with nothing to do
quiet-hourspersistent memory12 idle h per night14 nights
an agent with long-term memory got twelve idle hours a night and no objective. i sorted every action it took into categories and counted compute, not actions, so long operations weigh what they cost.
what i thinkleft alone, it mostly maintained itself. on night 9 it asked for a bigger memory store, and that's when i paused the run to think properly about what i was rewarding by saying yes.
exp-004done 19.06.26
the grader and the graded
self-written evals4 skills40 items eachheld-out control
the model writes a 40 item eval for a skill and then takes it. afterwards it sits a held-out set on the same skill, written by a different model and matched for difficulty by human raters.
its own evalheld-out eval
fig. 8 · pass rate, %gap sorted largest first
what i thinkthe gap is widest on the skills that are hardest to specify. it also reused phrasing from its own answers when it wrote the questions, which is a polite way of saying it wrote the answer key.