Delegated intelligence deployed by Otman Mechbal across Claude Code, Codex, Hermes, Grok build CLI and OpenCode, with OpenClaw as the retired predecessor. Each bar is read locally from that tool's own session log or database. The point is not the number. It is what the number moved.
9.95B
total tokens · 169 days operating · since 27 Jan 2026
596.8M
biggest day · 2026-06-02
71.1M
avg per active day
$12,635
equivalent value at API list price
Daily burn (2026-01-18 → 2026-07-15)
Each square is a day. Gold = heavier delegation. Empty = idle. Log scale, so a near-billion day and a few-million day can share one grid. 140 of these days logged token activity; earlier OpenClaw days predate the surviving logs.
Top 10 days
2026-06-02596.8M
2026-06-01573.5M
2026-06-06343.9M
2026-06-15336.5M
2026-05-18331.2M
2026-05-20319.5M
2026-05-21288.7M
2026-06-04285.1M
2026-07-14282.6M
2026-05-22280.9M
A spike only counts as proof once it is tied to what it moved. That mapping lives in the Build Log.
Adversarial by default (I do not trust one model)
Claude6.58B (66%)
Codex3.03B (30%)
Hermes275.3M (3%)
Grok build CLI51.1M (1%)
OpenClaw44.3M (0%)
OpenCode4.2M (0%)
Other (Gemini, Droid, etc.)0 (0%)
This split is the point, not the volume. I run models against each other: Claude for careful reasoning, Codex and Hermes for a second opinion and an adversarial check, Grok build CLI and OpenCode in the mix, OpenClaw as the retired predecessor to Hermes. Most people run one model and accept its answer. On high consequence work, one model's confidence is not evidence, so I triangulate. Each bar is read from that tool's own local log or database, not guessed from a model name, so the split itself is honest, not just the total. Claude Code prunes old session logs, so its number is held at the highest total ccusage ever measured for each day, its June 24 record plus everything burned since, never an amount I cannot show. Grok's number is a floor read from its surviving rolling log, not a complete lifetime count. "Other" is whatever is left after the named tools: mainly Gemini CLI and Droid, both real but small. OpenClaw total (44.3M): estimated from recovered OpenClaw logs (6.3M tokens, 12 active days, 2026-04-16 onward) plus confirmed daily operating history from Feb 3, 2026.
From OpenClaw to Hermes
My first delegation experiments ran on OpenClaw, starting late January 2026. It showed me what an always on agent could do, and what it cost: heavy to maintain, thin memory, quick to break. By late May I had moved the daily work to Hermes, which carries a stronger memory layer, improves its own behaviour over time, and asks far less upkeep. OpenClaw stays here as the honest predecessor. The early, messier curve is part of the record, not something I hide.
This is not a flex. Here is how to read it.
A token count is a trace, not a trophy. A big day can mean hard problems solved, or it can mean an agent searched the wrong folder and retried a broken command. Volume alone proves nothing. The number opens the door. What it moved is the proof, and that lives in the build log, not in this chart.
The metric that matters: value, not volume
Tokens are the calorie, not the nutrition. The real question for a business is not "how much AI did you use," it is "how much verified, useful work per dollar, per minute, per unit of risk." These are the metrics I instrument per workflow. The volume chart above is the hook. This is the substance.
Metric
What it measures
Cost per successful workflow
Real unit economics: AI cost plus human review plus rework, per accepted result
Time saved per verified deliverable
Manual baseline minus AI-assisted time. Real productivity, not activity.
Human review minutes per output
The hidden tax. Did AI remove work or just move it into checking?
Token cost per useful decision
Whether the reasoning was efficient or wandered
Blast radius if the agent fails
How much damage one wrong action can do
Reversibility
Can the action be undone? The safest automation fails inside a contained zone.
Evidence quality
Supported claims over total claims. Polished output with weak evidence is delayed damage.
Per-workflow instrumentation of these is the next build. Shown here as the model, not yet as live numbers, so nothing is faked. Most dashboards show usage. This one is being built to show value, waste, risk, and trust.
Cache read is ~89% of the total. That is the fingerprint of sustained agent loops carrying work across many steps, which is computer work, not chat. Always shown, never hidden inside the headline.