FactoryFloor

A framework for measuring the effectiveness of software factories.

What is the factory shipping?

01 / What shipped

Don’t count tickets and PRs. Track every new feature, the bugs that followed it, and the improvements made later so you know where your team’s time and tokens are going.

Insights

  • How are we balancing speed and quality?
  • How much of our time goes to new features vs. KTLO?
  • Is the maintenance cost of what we’ve shipped eroding our productivity gains?

Critical metrics

  • Human effort and token cost behind everything that ships, broken down by feature and work type (bug, KTLO, improvement).
  • Human effort to ship every feature and fix
  • Token cost per PR, work item, and team.

Fig 1.1Shipped this week

Sample data

  • Bulk export

    6 PRs · 2 follow-up bugs · 9.4 agent hr · 1.8 human hr · $212

  • Webhook delivery queue

    3 PRs · no follow-ups · 2.1 agent hr · 0.6 human hr · $48

  • New billing module

    4 PRs · 1 follow-up bug · 7.7 agent hr · 2.3 human hr · $174

  • …and 22 other capabilities this week
Fig 1.1Shipped this week
WorkNew featuresBug fixesImprovements
Merged PRs (48)42% · 20 PRs31% · 15 PRs27% · 13 PRs
Engineering hours (196 hr)42% · 82 hr20% · 38.5 hr39% · 75.5 hr
Agent spend ($1,240)42% · $52017% · $21041% · $510
Bulk export70%20%10%
Webhook delivery queue10%85%5%
New billing module15%10%75%

Fig 1.2Where the work went, 8 weeks

Sample data

  • New features
  • Improvements
  • Bug fixes
Fig 1.2Where the work went, 8 weeks
WeekNew features (hr)Improvements (hr)Bug fixes (hr)Bug fixes ($)
Week 1141622$290
Week 2161620$265
Week 3171619$240
Week 4181717$215
Week 5191715$190
Week 6201814$165
Week 7211812$145
Week 8222010$130

How autonomous is the factory?

02 / Autonomy

In an effective software factory, the path from prompt to production is straight. Fewer mid-session corrections, less rework, fewer review cycles — and code that lands in production.

Critical metrics

  • Session autonomy: the steering and interruptions required to get good code.
  • PR rework required before merging.
  • Amount of code discarded locally or in closed PRs.
  • Differences in autonomy between local sessions, background agents, and automations.

Fig 2.1Prompt → production

Sample data

Fig 2.1Prompt → production
RunHuman interventionsSteps
Low autonomy7clarify, steer, share docs, re-prompt, correct, review, rewrite
Medium autonomy4clarify, re-prompt, review, correct
One-shot1review

Where does human time go?

03 / Human effort

In a healthy factory, agents give humans leverage — the ability to do more. So where is their time going? What are the new bottlenecks? And how are the most effective developers working?

Critical metrics

  • Human effort and tokens for each PR.
  • Waiting time: hours waiting on review, or on long-running agents.
  • Time spent reworking PRs.
  • Time spent reviewing PRs.
  • Token spend per person, broken down by output.

Fig 3.1The human week

Sample data

Fig 3.1The human week
Hours in the loopMONTUEWEDTHUFRITotal
Waiting on agents or review2222210
Reviewing PRs212117
Prompting and Steering3334215
Reworking121138

Fig 3.2AI maturity, per contributor

Sample data

Fig 3.2AI maturity, per contributor
ContributorMaturity (1–5)Level
R. Okonkwo5Orchestrates
J. Lee5Orchestrates
M. Patel4Agentic
S. Rivera4Agentic
A. Chen4Agentic
T. Novak3Assisted edits
D. Haas3Assisted edits
K. Murphy2Q&A

What does a unit of shipped work cost?

04 / Cost & ROI

Account for every human hour and every dollar your factory spends. Find what to optimize, name the real waste, and show a clear ROI for the work each team is doing.

Critical metrics

  • Session yield by count and dollars: shipped to production, rework, discarded, closed PR.
  • Cost per work type, with trends.
  • Estimated agent hours from each model.
  • Average cost per PR type.
  • Dollars spent on code review.

Fig 4.1Session yield

Sample data

Fig 4.1Session yield
OutcomeSessionsShareCost
RESEARCH & BRAINSTORM1149%$1.3K
SHIPPED TO PRODUCTION51243%$18.4K
SHIPPED AFTER REWORK21818%$8.9K
DISCARDED21418%$2.9K
PR CLOSED UNMERGED14612%$1.2K

Fig 4.2Leverage & unit cost

Sample data

Est. human hours from AI

1,840 hr

+18%vs last quarter

Avg. cost per feature

$41.11

−24%vs last quarter

Avg. cost per bug fix

$16.11

−31%vs last quarter

Avg. cost per code review

$21

−8%vs last quarter

Fig 4.2Leverage & unit cost
MeasureCurrentChange vs last quarter
Est. human hours from AI1,840 hr+18%
Avg. cost per feature$41.11−24%
Avg. cost per bug fix$16.11−31%
Avg. cost per code review$21−8%
Spend only means something next to what it bought. Agent hours are reported per model and deliberately never converted to an hourly rate — so the return shows up as dollars per shipped thing, not as a price on anyone’s time.

Factory Maturity

Everyone is figuring out how AI fits their daily work. These are the stages we see individuals and teams move through.

Team

  1. 1
    Asking AI

    Developers reach for AI to answer questions, but are not running agentic flows.

  2. 2
    AI writes the code

    Developers use agent mode and most PR code is AI-written. The team largely operates the same as before.

  3. 3
    Central team tunes the agents

    A central team drives the agentic SDLC — building custom skills, agents, governance, and MCPs to make agents more effective.

  4. 4
    Every team tunes the agents

    Each team owns improving AI performance on their part of the codebase — the outer loop is everyone’s job.

  5. 5
    Prompt to production

    There is a straight line from prompt to production. ~50% of work is kicked off from Slack or Linear, accept rates are high, and low-risk PRs automerge.

IC

  1. 1
    Tab-completer

    Still working the way they did before — AI as autocomplete.

  2. 2
    Single agent

    AI-aided development, working on one PR at a time.

  3. 3
    Multi-agent

    Managing multiple agents and PRs at once, focusing on plans, tests, and specs to stay aligned with the agents.

  4. 4
    Harness engineer

    AI writes every PR. When it’s not working, goes up a level and adjusts the harness.

Track AI Code all the way to production

Observability for your software factory.

Tracks code from every major coding agent

  • Cursor
  • Codex
  • Claude Code
  • GitHub Copilot
  • Gemini
  • OpenCode
  • Junie
  • Droid
  • Windsurf
  • Amp