More agents are not automatically a better organization. Before making an engineering workflow parallel, I want to know what can truly be done independently.

My default preference is a serial loop with explicit planning, action, and review. These are responsibilities, not necessarily three permanently running agents. The right implementation may use different models or the same runtime at different moments.

Make a handoff an artifact

A useful plan states the objective, the allowed scope, the current evidence, and the next acceptance condition. It should be specific enough that an actor can proceed without reconstructing the planner’s entire conversation.

The actor then changes the engineering state and records what happened: modified artifacts, commands or experiments, results, failures, and unresolved questions. The reviewer compares those results with the objective and writes the next bounded step.

Plan → Act → Inspect evidence → Revise

The durable state lives outside any one model context. A restart should require reading the current handoff and evidence, not interpreting a terminal transcript as the only source of truth.

Spend capability where it changes the outcome

A stronger model may be most useful for framing the task, diagnosing a difficult failure, or reviewing a consequential change. A less expensive model may be sufficient for a clear implementation step. That is a hypothesis to test on the actual workload, not a universal prescription.

The comparison needs more than token prices. Rework, tool misuse, abandoned sessions, review effort, and simulation waste can outweigh an apparent saving. I would compare complete accepted outcomes under a shared task definition.

Protect the reviewer’s role

A reviewer that merely reads the actor’s summary is receiving the actor’s interpretation of the evidence. For consequential claims, it needs access to the artifacts and checks that support the summary.

Even then, model-based review is not the same as a deterministic verifier. A second model can identify omissions or suggest a better experiment, but it should not substitute its confidence for an electrical acceptance test, a file-identity check, or a budget boundary.

This is another reason to keep verification separate from orchestration. The same formal acceptance path should apply regardless of which model planned or executed the work.

Parallelize work, not ambiguity

Parallelism is attractive when tasks have stable interfaces and independent state: evaluating separate candidates, inspecting different source areas, or checking distinct hypotheses. It is less attractive when two actors are modifying the same evolving design and neither knows which version the other has seen.

I do not oppose multiple agents. I oppose using concurrency to hide an unresolved coordination problem. Before parallelizing, define ownership, merge conditions, evidence identity, and what happens when the branches disagree.

Prefer a small, recoverable loop

The orchestrator should do the unglamorous work reliably: preserve state, detect completion, enforce budgets, and resume from a known boundary. It should not become a second designer made of ever-growing rules that infer the next RF strategy.

A well-designed serial loop is easier to inspect and a useful baseline against which to measure a more elaborate organization. Add complexity when it solves an observed limitation—not because a diagram with more agents looks more capable.

The objective is not to keep every model busy. It is to move the engineering artifact forward, with less uncertainty about what was changed and why.

Ideas in progress. Corrections welcome.

Find me online