Reviewing a week of agent transcripts to revise skills
What 15 sessions and 132 worker transcripts showed about my written procedures, and the four new skills, four scripts and one hook warning that came out of it.
I had an agent check Blobber’s skills against a week of the sessions’ own transcripts, and then I had all seven of its proposals built. A skill here is a written procedure that an agent session loads for one kind of task, like landing a branch or briefing a worker. The review found that the sessions kept working out the same four or five procedures by hand, and that workers were spending most of their context reading before they built anything. Nothing in the game changed on 10 October because of this.
Reviewing the transcripts with counts instead of reading them
The review covered 15 main sessions and 132 worker transcripts. It didn’t read them through. Extraction scripts pulled out five things: my own messages, the timeline of tool calls, statistics for each worker, the opening of each worker’s final report, and sentences that repeat across briefs. The scripts and extracts were deleted afterwards, and nothing in the repository was touched.
The result was one note in four parts. The first part goes through each existing skill and lists what the sessions did that the skill doesn’t say, and what the skill says that nobody follows. The second lists new skills, ranked by how many calls they’d save. The third lists things that should be a hook, a script or a standing rule and not a skill. The fourth says what it looked for and didn’t find, and what it didn’t look at.
That last part threw out two ideas. A skill for tracing a flaky test had only three cases in the week, which isn’t enough. A skill for studying picture sheets and films before showing them wasn’t needed, because the sessions already did that the way the existing skill says.
What were the sessions doing by hand?
Landing a branch was the clearest case. The landing skill covered about 5 calls, and a real landing was 9 to 28. Every landing went on through four things the skill didn’t name: the shared work-list page, a devlog entry, a patch to the handoff document, and sending the film again. The main session wrote 28 devlog entries by hand in two days, at 5 to 9 KB each.
Starting a session took 15 to 44 calls, and it was the same eight steps every time. Two or three of those were small inline filters for “answered and not done yet” on the work-list page, rewritten in each session.
The work-list page took 208 calls to its data tool in the week and 67 patch scripts. Two batch writes failed validation because of a stray key.
The handoff document got asked for 8 times, three of them in one hour. It grew from 20 KB to 32 KB in one day, and every new session reads it whole. The standing instructions said to overwrite it, and the sessions patched it.
Four landings in one day had a merge conflict in the work list or the tests’ README. Each one was the “both sides added lines” kind, and each was settled by a new throwaway Python script.
I also couldn’t answer a lot of the questions I was asked. At least eleven times in three days I replied with something like “I don’t understand the question” or “what are the tradeoffs”, because the question pointed at a report I didn’t have in front of me.
Why were workers reading and not building?
After the context checkpoint hook went in, 13 of 27 worker reports were handovers to a fresh worker, and one task went through 5 workers in a row. Five first workers got to 190,000 tokens or more before they wrote their first project file, or never wrote one. They’d read whole source files of 41 to 51 KB and two skill files of 40 and 27 KB. Every worker also starts at about 40,000 tokens before its first call.
The two modeling skills were part of the load. The pipeline document alone was read by 20 of 62 workers, which is 454 KB of text. A worker building a new generator spent about 95 KB on skill files before it opened the kit.
Briefs were padded too. The skill for briefing a worker says not to repeat the standing instructions, but 13 of 62 briefs repeated the rule about whole-file reads and 15 opened with the same sentence. Workers already have all of that in context.
Turning each repeated procedure into a skill and a script
There are four new skills, and each one is 3 to 4 KB.
start-session is four calls and one reply. A script, tools/session_start.py, prints the state of the main branch, every worktree with its uncommitted count and whether its note says it handed over, the model workbench’s picks, the first section of the handoff and the open decisions. It can also filter saved page data for answers I’ve typed.
worklist-page has the page’s data layout and the exact shape of a batch entry. tools/worklist_page.py shows one item cut to a length, gives the next free id, appends to an item, ticks it, and writes the page documents for new items.
devlog-entry says where an entry goes, in what order to publish, and that the site is public. It also says the entry is a worker’s job, written from a brief that names where the facts are.
session-handoff fixes the handoff’s headings, says it gets overwritten whole and never patched, and caps it at 6 KB.
The landing skill now has 12 steps and runs to the real end of a landing. It also says a line in the handoff isn’t me asking for a merge. Twice a new session had tried to land a branch because the handoff said to, and the permission check refused both.
I had the two modeling skills slimmed by moving material into named files next to them, with nothing deleted. One went from 38.5 to 16.4 KB and the other from 31.5 to 12.9 KB. A new script, tools/check_skills.py, takes everything in backticks in the skills and the standing instructions, tests whatever looks like a path, and fails if one doesn’t exist. It checks 294 paths.
The standing instructions got a section on asking me a question. Each question has to make sense without the report, state what each answer costs, give a recommendation and say what happens if I don’t answer.
Keeping both sides of a merge conflict, unless an id repeats
tools/keep_both.py replaces the throwaway resolvers. Every conflict hunk becomes the main branch’s lines followed by the incoming branch’s, and a line that both sides added word for word is kept once. Then it looks for an id that’s now on two lines. An id is a work-list item or the first cell of a table row:
IDS = (re.compile(r"^- \[[ x]\] \*\*([A-Z]+[0-9]+[a-z]?) "), re.compile(r"^\| `([^`]+)` \|"))
A repeated id means both branches changed the same item, and keeping both doesn’t settle that. In that case the script writes nothing, prints the ids with their line numbers and exits 1. That check is why I didn’t use git’s merge=union attribute for those files. Union keeps both versions of the item and doesn’t tell you.
Warning a worker that hasn’t written anything by 120,000 tokens
The rule for workers is now that the first project file gets written by 120,000 tokens of context, or the worker stops and reports what it still needs to know. The context checkpoint hook warns about it once. It never blocks a call.
READ_AT = int(os.environ.get("CHECKPOINT_READING", 120000)) # a worker with no project write by here is warned
BUILDS = ("Write", "Edit", "MultiEdit", "NotebookEdit")
On each tool call from a worker, the hook looks for a small state file for that agent. If the file exists, the worker has either built something or been warned already, so the hook does nothing and doesn’t read the transcript. If the call is one of BUILDS on a path outside the temp folder, the hook records that the worker has built. Otherwise it reads the context size from the tail of the worker’s transcript, and at 120,000 or more it records the warning and sends it back as added context. A write to the worker’s own running note doesn’t count, because the note is in the temp folder.
The warning tells the worker to write the first file now, rough, or to stop and say exactly which file, function or decision it still needs. A report like that gets answered with a better brief. Briefs now give line ranges or function names for any large file, name the sections of a skill to read, and are saved to a file per branch so a carry-on worker can get the original word for word.
Refusing a long shell heredoc in a PreToolUse hook
The standing instructions already said not to use a shell heredoc for anything over a few lines or anything with a backslash in it. On this Windows machine large heredocs fail, and a doubled backslash gets silently collapsed to one. The review counted 174 python - <<'EOF' blocks in the main sessions anyway, and one worker’s note had lost its backslashes that way. The worker doing the review broke the rule three times itself, and each one cost a repair.
The choice was to enforce the rule or drop it. Enforcing means a refused command interrupts sessions until the habit goes, and I picked enforcing. The hook that already refuses a second whole read of a big file now also refuses a Bash heredoc with more than 5 lines or a backslash in its body. It applies the same two tests to a PowerShell here-string, unless the here-string is the message of git commit -m, because that’s the PowerShell tool’s own documented way to pass one.
The hook gets the tool call as JSON and matches the command text against these patterns:
HEREDOC_LINES = 5
HEREDOC = re.compile(r"(?<!<)<<(?!<)-?[ \t]*(['\"]?)([A-Za-z_][A-Za-z0-9_]*)\1[^\n]*\n(.*?)(?:\n[ \t]*\2[ \t]*(?:\n|$)|$)", re.S)
HERESTRING = re.compile(r"@(['\"])[ \t]*\r?\n(.*?)(?:\r?\n)?\1@", re.S)
GIT_COMMIT_M = re.compile(r"git\s+commit\b[^\n@]*-m\s+$")
HEREDOC wants a << with no third < on either side, so Bash’s <<< isn’t matched. Then it takes an optional -, the delimiter word with or without quotes, and the body up to a line that has only that word, or up to the end of the command. It searches the whole command, so a heredoc in the middle of a cd x && python - <<'EOF' chain is found too. For PowerShell, GIT_COMMIT_M is tried on the text before each here-string, and a match skips it.
The two tests are short:
for body, word in found:
count = len(body.splitlines())
if count > HEREDOC_LINES or "\\" in body:
So 5 lines pass, 6 are refused, and one backslash in a two-line heredoc is refused. The refusal names the delimiter word and which test failed, and then says what to do instead: write the script with the Write tool into the scratch folder, run it, delete it. The whole check is wrapped in a try, and any exception lets the command through.
The self-test went from 10 checks to 26 and all 26 pass. The hook only reads the command’s text, so a heredoc body made by command substitution isn’t parsed.
What hasn’t been tested?
None of the new skills had been used in a real session when this merged. There have been two uses since. The work-log entry behind this post was written by a worker following devlog-entry. And the main session reported that the merge that brought the skills in had one conflict, and it settled it with keep_both.py.
The hook’s warning passed its self-test (19 of 19) and was fired by hand, but I haven’t seen it in a live worker. The two page queries in start-session were written from the tool’s schema and haven’t been run.
The heredoc check has only been run by its self-test. I haven’t seen it refuse a command in a live session yet.
I didn’t split the pipeline document into its list and its rules, because a self-test reads four tables from that path.