AI Coding Factories

A coding agent can edit files, run tests, and open a pull request. Turning those capabilities into a repeatable delivery process requires a system that defines the work, supplies context, coordinates execution, verifies results, and learns from production.

That is what I mean by an AI coding factory: an engineering workflow in which agents carry work through defined stages, with explicit criteria for advancing, retrying, or asking for help. You can build one with existing models and coding tools, without owning the infrastructure that trains them.

AI coding factory

Paul Iusztin's Inside a Software Factory is a useful reference. He describes a lifecycle spanning intake, planning, implementation, review, release, and production feedback, supported by shared context, and he also documents an overbuilt first version he stopped using, which is a good reminder that the process has to earn its complexity.

Delivery loop

Request β†’ specification β†’ implementation β†’ verification β†’ release β†’ production feedback

Each transition needs an artifact and a decision:

  • A specification defines expected behavior.
  • Implementation produces a reviewable change.
  • Verification produces evidence that the change meets the specification.
  • Release produces an observable change in the running product.
  • Production feedback becomes the next request.

Factory architecture

I would organize the architecture around six components:

  • Context and specifications. Repository instructions, product requirements, architecture decisions, examples, and acceptance criteria give agents a working understanding of the system. Keep this material versioned and close to the code, with current decisions separated from historical notes.
  • Models and harnesses. The model reasons about the task while the harness manages file access, command execution, browser use, permissions, and context, so evaluate model choice and harness configuration together.
  • Orchestration and state. A coordinator tracks tasks, dependencies, assignments, retries, budgets, and completion. Durable records should make it possible to understand what happened and continue after an interruption.
  • Execution environments. Reproducible sandboxes hold dependencies and control access to files, credentials, and networks. Branches and worktrees isolate code changes, and containers or virtual machines bound execution.
  • Verification and release gates. Tests, type checks, security checks, browser validation, and review determine whether a change meets its specification. Merging and deploying should have explicit conditions.
  • Observability and feedback. Logs, artifacts, costs, failures, and production signals let engineers diagnose runs and improve the workflow.

Useful autonomy still depends on the configured system around the model, the same architectural concern as AI Agents are Infrastructure.

Orchestration and evidence

Anthropic's guidance on effective agents distinguishes predefined workflows from agents that dynamically choose their steps. A factory can combine both: a sequential pipeline for predictable work, parallel workers for independent tasks, and a dynamic split when subtasks emerge during investigation.

Parallelism needs boundaries. Two unrelated bug fixes can run together, while a database migration and a feature that depends on its schema need sequencing or an agreed contract, so I would define dependencies and file ownership before adding workers.

Long-running work needs continuity. Anthropic's research on long-running harnesses uses incremental progress and explicit handoff artifacts to preserve work across sessions. For a coding factory, that means recording completed tasks, failed checks, unresolved decisions, and the next useful step.

My practical rule is observable evidence: a regression test reproduces the bug, the fix passes it, existing behavior remains intact, and the relevant user flow works. Agent review can supplement those checks. Keep acceptance criteria stable while the implementation iterates.

Platforms and pieces

I think Factory is helping lead the way because it treats these concerns as parts of one delivery system. Its Software Factory documentation covers triage, implementation, validation, release, documentation, and monitoring, with visibility into automation health and repository coverage. The dedicated Software Factory product is currently in private preview.

You can also assemble a factory around existing tools. These projects address complementary layers:

  • Sandcastle is a TypeScript library for orchestrating coding agents in sandboxes, with configurable branch strategies and providers including Docker, Podman, and Vercel.
  • Machinist is open-source infrastructure for repeatable agent execution, approved commands, local repository access, and inspectable run records, and it leaves the shipping decision with a human. It is early-access software, so workflows must own any checkpointing they require.
  • Superpowers encodes a development methodology as composable agent skills. It runs brainstorming, planning, isolated worktrees, subagent-driven implementation, test-driven development, and review as mandatory workflows inside existing coding harnesses.
  • Addy Osmani's agent skills encode engineering workflows for specifications, incremental implementation, testing, review, and shipping, and carry that process knowledge into an existing agent environment.

Skills describe how to perform work, orchestration determines when it runs, and execution infrastructure supplies the environment and records the outcome, which leaves the connections between them as engineering decisions you still have to make.

For a small team, I would begin with a coding harness, a few carefully chosen skills, an isolated environment, and existing CI. A managed platform becomes more attractive when coordinating repositories, users, integrations, and operational controls takes substantial time.

Start with one task

The engineer still owns the problem, architectural tradeoffs, and definition of done, as in Engineering in the AI Era. People also need to see what is running, why it stopped, and how to redirect it, as in Evolution of AI UX.

I would start with one recurring task and measure time to an accepted change, review effort, defects, and total cost, use failed runs to improve specifications, tools, tests, and handoffs, and expand automation only when that evidence shows the factory is making delivery more reliable.

Related writing