Birdoggydog's Builds

Measuring agent token use in transcripts and guarding reads

A report over Claude Code's own session files, two measured constants, a hook that refuses a second whole read, and a handover at 250,000 tokens of context.

Blobber now has a script that reads the coding assistant’s own transcripts and reports where the tokens went. Two hooks came out of what it found: one refuses a second whole read of a large file, and one tells an agent whose context has grown past a line to write up its note and stop. Reading whole files is what fills a context here, not the number of tool calls. This was 10 October, and parts of it are still being switched on.

Reading the transcripts for where the tokens went

tools/usage_report.py reports by session and by background worker, by run of fifty requests, by tool, and the single tool results that cost most to keep in context.

I’d read a comment by someone who runs several agent sessions at once and hands each agent over to a fresh one at 45 to 50 tool calls. I wanted automated ways to constrain token use, built in as hooks. Nobody knew where this project’s tokens went, so the report came first.

The transcripts are one JSON-lines file a session, with one more for each worker in a subagents folder beside it. They sit under .claude/projects in the home folder, in a folder named after the checkout’s path with every character that is not a letter or digit turned into a hyphen. A worker’s label comes from agent-<id>.meta.json beside its transcript, which holds the description it was launched with.

For each transcript:

  • Rows of type assistant carry message.usage, so cost is added up without an estimate. One message is written once for each content block, so rows are keyed by message.id and the last usage for an id is kept. Without that, every request with two tool calls is counted twice.
  • A request’s context is input_tokens + cache_creation_input_tokens + cache_read_input_tokens. The peak is the largest of these.
  • tool_use blocks give each call’s tool and a one-line description of its input. tool_result blocks in the following user row are matched to them by tool_use_id and sized in characters, with pictures counted apart.

A weighted cost puts the four kinds of token on one scale at the published price ratios, in units of one fresh input token:

WEIGHT = {"fresh": 1.0, "write_1h": 2.0, "write_5m": 1.25, "read": 0.1, "out": 5.0}

The one-hour and five-minute cache writes are told apart by usage.cache_creation.ephemeral_1h_input_tokens and ephemeral_5m_input_tokens. The total is a ratio for comparing transcripts, not a bill.

The figure that finds a wasteful call is carried cost: a result’s estimated tokens times the number of requests made after it in the same transcript. A 12,000 token file read at request 10 of 130 is paid for 120 times.

Why didn’t four characters a token add up?

The first run sized tool results with the usual guess of four characters a token, and 1,500 tokens a picture. That put about 114,000 tokens of results in a transcript whose context peaked at 400,000 to 730,000.

So both constants were measured. Between one request and the next, the context grows by the previous request’s output plus whatever tool results came back. Turns that returned only pictures (under 400 characters of text) give tokens a picture. Turns that returned only text of more than 2,000 characters give characters a token.

Over 302 picture turns the median was 2,671 tokens a picture, with quartiles of 1,918 and 3,530. Over 787 text turns the median was 1.71 characters a token. That is far from four because the growth includes the reminders the tool adds beside a result, and because file reads come back with a line number on every line. With the measured constants the sum matches the peaks.

A worker’s first request was 38,992 tokens at the median: the standing instructions, the skills list and the tool definitions.

Would a handover every fifty calls save anything?

Not much here. Over 35 transcripts from 9 and 10 October (5 sessions and 30 workers), a worker’s first fifty requests cost about 1.6 million weighted tokens and each later fifty about 2.1 million. The context fills early and then stays large.

RequestsWorkers that got thereWeighted tokens, all workersA worker
1 to 503046.6 million1.6 million
51 to 1002246.8 million2.1 million
101 to 150714.4 million2.1 million

The weighted total was 128.2 million, 84% of it in workers. Input read back from the cache was 804.5 million tokens and 63% of the weighted cost. Output was 885,000 tokens and 3%.

What fills a context is whole-file reads. Read was 1,137 calls and 5.3 million tokens of results, which cost 348.8 million to carry. Bash was 2,013 calls, 3.3 million tokens and 188.5 million carried. Every other tool was under 3 million carried. Most of the 40 dearest results were whole reads of the model generator’s kit files, several read twice by one worker.

So the order of work changed: a cheaper model for mechanical workers, a guard on repeated whole reads, a cap on command output, and a checkpoint by context size and not by call count.

Refusing a second whole read of a large file

tools/hooks/token_guard.py is a PreToolUse hook for Read, Bash and PowerShell. It gets the call as JSON on standard input and answers on standard output.

A Read with an offset, a limit or pages is always allowed. So are pictures, documents and models, by extension. Otherwise the file’s lines are counted, and under 400 lines it is allowed. For a longer file the hook keeps a small JSON file for each session and agent, listing the paths read whole. The first whole read is recorded and allowed. The second is refused.

A shell command that runs grep, rg or cat on the work list is also refused, unless something in the pipeline shortens the lines (cut, wc, awk, a count or names-only flag). That file’s items are single lines of several thousand characters.

A refusal is this, printed with exit code 0:

{"hookSpecificOutput": {"hookEventName": "PreToolUse", "permissionDecision": "deny",
  "permissionDecisionReason": "token guard: you have already read body_kit.py whole ..."}}

The reason says what to do instead: search for the name, then read a range. The hook fails open. Any exception, or input that is not JSON, and the call goes through.

The standing instructions already said not to dump large files. Several workers still read the same 900-line kit file whole twice, in part because another rule told them to read a file again after changing it from the shell. An instruction is read once at the start of a long context. A hook answers at the moment of the call and names the cheaper form. That other rule now says to read only the changed range.

Handing over at a context size instead of compacting

tools/hooks/context_checkpoint.py is a PostToolUse hook with no matcher. It reads the last 400,000 bytes of the agent’s own transcript, walks the lines backwards to the first one holding "usage", and adds the three input figures. That is the context of the agent’s last request, with no parsing of a file that can be 25 MB.

The lines are 250,000 tokens for a worker and 300,000 for the main session. Both are guesses. The hook speaks once on crossing the line and again each 50,000 tokens after.

What it says goes back to the model as hookSpecificOutput.additionalContext. A worker is told to finish the one step in hand and bring its running note up to date: done, left to do in order, verified, in flight, and what the next worker must know, with paths and line numbers. Then it stops with a report whose first line is exactly HANDOFF: cut at the context checkpoint, resume from the note.

The first proposal was the tool’s own compaction at a size. I turned that down and kept the handoff model. My experience of compaction is that the workers get confused and don’t have as good a sense of what’s happening as they do after a handoff made on purpose. Nothing is compacted.

On a HANDOFF: report the main session reads the note and asks the same worker for anything missing while it still holds the context. Then it starts a fresh worker in the same worktree, with the original brief word for word and an instruction to read the note first.

Why did a worker hit the line before building anything?

Both hooks fired inside a worker the same day. The first worker on a model task, themes and gear for the new orc, reached 255,461 tokens of context in 22 requests and under five minutes. It had read kit files whole and built nothing. A second worker, told to read ranges only, used the same allowance to write three kit modules and build the themed orc to step 11 of the pipeline before it handed over too.

Two changes followed. tools/blender/kit_map.py walks every .py under the Blender tools folder and parses each with Python’s ast, so nothing is imported or run. It writes map/INDEX.md, one line a file, and map/<file>.md, one line for each function, class, method and multi-line table of capitals, with its line range, its arguments and the first sentence of its docstring. A worker reads the map and then a range. A --check mode fails if any map on disk is stale, and a test in the game’s suite runs it.

76 source files are mapped in 125 KB. The largest kit file is 101 KB and its map is 6 KB. The dozen files the orc task turns on are 572 KB.

A brief for kit work may also carry one line, CONTEXT CHECKPOINT FOR THIS TASK: <number>, that moves that worker’s line, held between 250,000 and 400,000. The hook reads it from the first 200,000 bytes of the worker’s own transcript, where its brief is. No setting or environment variable is involved, so two workers running at once can have different lines.

Also done

  • The skill that briefs workers now picks a model. The smaller one is for a brief that says exactly what to do, with a check that says whether it was done. The largest is for a study, a new generator, anything judged by eye or an untraced fault. In doubt it’s the largest, because a wrong build costs more than the difference.
  • The cap on command output needed no hook. The tool has a setting, bashOutputMaxChars, with a default of 30,000. It is going to 10,000.
  • One mistake: the skill that documents the settings file was loaded to check the hook format, and it printed its whole schema, tens of thousands of tokens, into the session that was trying to save them.