11 October 2026

A week of running Orka ​

On Sunday I asked an agent to skim every coding session on my machine from this week. I drove 46 of them myself. In the background, Orka ran 589 worker runs across four projects. Most of the code this week was written by models working through cards I wrote, and reviewed by other models before I read it.

Orka is the open-source skill I built for exactly this. One agent acts as the CTO: it plans, writes task cards, sends each card to a coding CLI in its own git worktree, has a different model review the diff, and records how every run went in a ledger. Every number below comes from that ledger and from the runs’ own metadata, not from my memory.

The week in numbers ​

Worker runs589, of which 575 exited cleanly
Projects4
Work cards111, plus their reviews
Commits from workers882
Worker timeabout 125 hours, in six days
Median runa little over 8 minutes
Busiest dayFriday, 153 runs

The four projects were a product catalogue system (63 cards, including QA and fix cards), a backup organiser (22), an image-processing tool (18) and a fitness app (8). Monday had no runs at all; that was planning. Sunday had 21, mostly the last fixes before a release.

Who did the work ​

Every card goes to a role, and every role to a model:

  • Lead, 258 runs. Reviews and critiques, nearly all by GPT-6 Astra. This is the most-run role by far.
  • Senior, 223 runs. The hard cards, split mostly between Claude Opus sub-agents and Codex.
  • Mid, 75 runs and junior, 33 runs. Well-specified cards and browser QA, mostly Sonnet and Grok.

By CLI, Codex did 307 runs, Claude sub-agents 216 and Grok 63.

The reviews are the point ​

There were 132 first-round reviews. 118 of them came back REQUEST CHANGES. Only 14 approved on the first try.

These weren’t nitpicks. The worst ones from this week’s ledger:

  • A plain save that silently deleted hidden and archived values. Data loss, with tests passing.
  • GET handlers that changed data: a CSRF hole.
  • Malformed import rows that set stock to zero. That was the exact risk the card had warned about.
  • Duplicate options in a product configuration that created stock out of nothing.
  • An integration that made up the shape of a third-party API response and rounded money to zero.

Every one of these got past the author’s own tests. A second model, reading the diff cold, caught them. That gap is why Orka exists.

What the ledger says about the models ​

Before every run the orchestrator predicts four scores, and after it records what actually happened: smart (did it crack the hard part), dumb (how bad was its worst mistake, lower is better), speed and cost. These are this week’s averages over 262 recorded runs:

WorkerRunsSmartDumb ↓SpeedCost
Codex · GPT-6 Astra · high1049268078
Grok 4.7 · high138457583
Claude Opus sub-agent4083235256
Claude Sonnet sub-agent2482197079
Grok CLI · 4.7 · high2777166084
Codex · GPT-6 Sol · high3174236660

Some readings:

  • Astra is the most reliable reviewer I have. That’s why it ended up doing almost every lead review.
  • Opus clears hard cards but is the slowest and most expensive here, and its worst mistakes are not small.
  • Grok 4.7 at high effort is strong and cheap. The same model through its own CLI is less steady, and it was behind the worst single moment of the week: during browser QA it got stuck in a loop, printed about 8,500 identical lines, and left the server and the browser running.
  • The orchestrator’s predictions were mostly cautious. Over 146 cards, the actual smart score beat the prediction by more than ten points 43 times and fell short of it by that much only 13 times.

What the ledger says about me ​

At the end of each project the orchestrator writes a retro. Its own mistakes go in the ledger too. This week’s were not flattering:

  • Oversized cards. On the biggest project, an independent retro counted about 63 wasted attempts. One card took 11 tries, another 9. Almost every time the cause was a card that was too big, or one that left out a security or concurrency requirement.
  • Merging on thin evidence. Two cards were merged with QA that was too weak, and at least once the review fixes were never checked against the final commit.
  • Scope creep. A cosmetic fix grew into a change to an approval contract.
  • Waiting for me. Too many fixes happened only after I asked for them.

The lesson I keep relearning: throughput isn’t the hard part anymore. 589 runs in a week isn’t hard work, it’s a queue. The hard parts are a card small and clear enough that the agent can’t misread it, and a review strict enough to catch it when the agent misreads it anyway. The process rules that came out of this week’s retros are short: spec, then card, then an evidence checklist; redesign if the same kind of finding comes back twice; watch a worker’s progress by its file times, not its narration.

The rest of the week ​

The sessions I drove myself were the ones that needed judgement, not output: model choice and memory design for an AI assistant platform, a narrowly scoped API token for email routing, a model-router update, and changes to Orka itself.

The homelab came up every day: the maintenance day and what followed it, a second NVMe drive, some DNS trouble, a reinstalled house assistant and a read through the UPS logs from a power cut.

And eight of my sessions this week said nothing but “reply with exactly: ok”. Those were health checks. Not everything an agent does is work.