MIT · v0.1.0 · alpha

Meet Fabrika:

An opinionated software factory -
Production-ready code,
optimized for human decisions.

You define the feature and rule on what comes back. In between, model agents read the repository, write the spec that you review, build, test and repair as needed. What comes back is a review packet: contextualized diffs, plus evidence that the code does what you specified.

Feature: Card Tags Project: Kanban Board

  1. repo ready
  2. intent
  3. questions
  4. spec review
  5. build
  6. review
  7. feature review · you

Plan

  1. 01 scout
  2. 02 interrogator
  3. 03 spec writer

Build

  1. 04 architect
  2. 05 workers
  3. 06 integrator

Verify

  1. 07 oracle
  2. 08 checks
  3. 09 breaker

Review

  1. 10 attribution
  2. 11 qa
  3. 12 review
  4. 13 arbiter
  5. 14 repairers
  6. 15 simplifier
  7. 16 rapporteur

16 plan

15 actual

2 called for help

passed called for help at work now nothing to do

Ships, with rulings

Tags on cards works as specified, but hold it for two rulings on what happens when the same tag is changed twice quickly.

Read how a packet is built

Why Fabrika

Fabrika vs
Your Favorite Coding Harness

A harness can do every step. A factory makes sure it does.

Your harness can spin up subagents, write tests and check its own work. Whether it will, and whether it will do it the same way twice, is up to the session.

Fabrika fixes the line instead. The flow is prescribed rather than configurable, which is what lets it guarantee red-teaming, blind tests and the project’s own checks on every feature, with plain code deciding what runs next.

Builders and checkers, on different model lineages.

Every station is an agent with one job and one prompt. Builders make things; checkers attack what was made. Fabrika wants the two teams on different model families: for example, build with Claude, check with GPT or Gemini.

The crew: who stands at each station, what model they run on, and where the line loops back for repairs.

Each agent sees what its job needs

A worker gets the whole branch and every tool the harness has. The oracle gets the requirements and none of the code.

Models matched to the job

A scout that reads and summarizes code can run on something like Haiku; the architect that plans the build needs something like Opus.

Prompts you can read and change

Each role is a Markdown prompt loaded at call time, tuned for its station and editable in the console.

Repo ready: Approve the checks

Point Fabrika at a repository and it surveys how the code is already checked, tested and guided: linters and type checkers, security and accessibility checks, unit, integration and user-facing tests, and agent guides like AGENTS.md and CLAUDE.md.

Every check it finds runs on an untouched checkout. Nothing is built until you accept them, because every later feature is judged against that baseline. A check does not have to be green to be accepted: fix it, hold it at today’s number, or rule that it does not apply.

A project once it is repo ready: checks, tests, guides, environment and what the runs cost.

Spec review: Freeze the spec

Give it a prompt and a name. The scout reads the codebase, and the interrogator asks what reading could not settle. Then a spec writer drafts and a spec checker scrutinizes; only the disagreements they cannot settle reach you, as open questions.

Every spec follows one outline, so you know where to look:

  1. 06 How this was decided

Each criterion says how it will be verified, at which level, and how hard it would be to undo, from easily changed to cannot be undone . The oracle writes its tests against exactly these. Nothing is built against a spec you have not frozen.

Before the spec: questions
What it thinks you asked for, and the questions it needed answered first.

The build:
watch,
don’t steer.

The architect plans the units of work and a plan checker objects; you see the plan, at plan review, only if they cannot agree. After that there is no input mid-run.

Two lanes run blind to each other: workers build while the oracle writes tests from the spec. Running the checks, running the oracle’s tests and splitting repairs into non-overlapping units is done by plain code, not by an agent.

The line as it ran: one lamp per station, the rework rail, and a timeline of every station.
The stations after the spec is frozen
Station Team What it does
worker builder Writes the code and its unit tests. Several run in parallel, each in its own unit of the plan.
integrator builder Joins the workers’ units into one whole, and writes the tests that cover the seams between them.
oracle checker Writes acceptance tests from the spec alone, in parallel with the workers. It never sees the implementation.
breaker checker Attacks the result: boundaries, failure paths, concurrency, hostile input, criteria met in letter but not in spirit. It writes probes that fail to prove it.
reviewer checker Makes the strongest case against shipping: duplication, broken conventions, unearned abstractions.
arbiter neither Rules on every finding: repair it, simplify it, send it back to the oracle, rule it out of scope, or escalate it to you. It cannot make a finding disappear.
repairer builder Makes the narrowest fix for what the arbiter sent back, in a loop bounded by rounds, budget and wall clock.
simplifier builder Simplifies code that is correct but more complicated than it needs to be. Everything is re-tested afterwards.
rapporteur neither Assembles the review packet.

Feature review: Rule on the packet

What comes back is not a pull request. It is a review packet: the code, plus a structured argument that it does what was specified. Code is reached through the criterion that asked for it. You audit the argument and rule on a handful of decisions, instead of reading a diff from the top.

The outcome is either a branch you can merge, or another round of work.

The decisions the run could not settle on its own: findings that were escalated or could not be repaired, and anything you need to check by hand. Each one carries the arbiter’s reasoning and the evidence behind it, and you repair or dismiss it with a reason.

Your calls: the three calls this run left for a person, a license blocker and two major findings, each with the arbiter’s reasoning and the evidence beside it.

As-built:
the codebase,
read file by file.

You need a mental model of a system before you change it. A reader agent reads every file, plain code checks what it reported and joins it into a graph, and a cartographer names the groups. The result is an architecture, a set of subsystems and a map of capabilities, each with a stable three-letter id.

It runs whenever you ask, and afterwards reads only the files that changed since the last commit it processed.

The as-built for a sample project: what it is, how it fits together, and what was read.

Three skins,
one board.

The console comes in default, day and night, switched from the masthead and remembered by your browser. A skin moves one ground onto the other rather than inverting the page: in day the board joins the floor, in night the floor joins the board.

Pick a skin, or drag across the screen.

The crew page in the default skin: enamel board and rail over a pale green-grey floor.

default The enamel board and its rail hang over a painted floor; the record is read on the floor.

Runs on the subscriptions you already have.

Fabrika drives Claude Code and Codex headlessly, on your Anthropic and OpenAI subscriptions. Anything OpenAI-compatible, OpenRouter or a model on your own machine, runs through the OpenHands harness.

Subscriptions get throttled, so each team has a fallback agent the factory switches to when its main route says no. Every model name lives in one file, factory.yaml.

Accounts that serve the models and bill for them; harnesses that run the agents.

What it is for, and what it is not.

Built for

  • One person, or a small team, building features largely on their own.
  • Well-defined, self-contained features, designed before they reach the factory.
  • Running on your own machine, as a single user.

Not yet, or not at all

  • Team logins, coordinating work streams, cloud deployment.
  • Work that needs a lot of back and forth, like shaping a user experience or exploring a problem space. Use an interactive agent session for that.

Where it stands

Alpha, version 0.1.0, MIT licensed. It has not been used in anger yet, and it will have bugs. About twelve hundred tests each assert a property that makes it worth running: the invariants, isolation, the repair loop, what the console draws.

Put a factory beside your repo.

  • Python 3.11+ and git
  • Docker: every check and agent runs in a container, with no fallback to your machine
  • A model provider: an OpenRouter key, any OpenAI-compatible API, or a Claude Code or Codex subscription
quick start
git clone https://github.com/AlexTatiyants/fabrika.gitcd fabrikapython3 -m venv .venv && source .venv/bin/activatepip install -e ".[dev]"cp factory.example.yaml factory.yamlexport OPENROUTER_API_KEY=sk-...uvicorn factory.server:app --port 8077

Then open 127.0.0.1:8077. On a fresh machine it opens on setup: it checks Docker and the way out, connects accounts, builds each harness and staffs the crew.