FactoryFloor
A framework for measuring the effectiveness of software factories.
What value ships?
What reaches customers — and is it valuable, or buggy?
02How autonomous is the factory?
How straight is the path from prompt to production?
03What are the humans doing?
Is AI real leverage — or just new bottlenecks?
04What does it cost?
What do tokens and human hours buy? Where’s the waste?
What is the factory shipping?
01 / What shipped
Don’t count tickets and PRs. Track every new feature, the bugs that followed it, and the improvements made later so you know where your team’s time and tokens are going.
Insights
- How are we balancing speed and quality?
- How much of our time goes to new features vs. KTLO?
- Is the maintenance cost of what we’ve shipped eroding our productivity gains?
Critical metrics
- Human effort and token cost behind everything that ships, broken down by feature and work type (bug, KTLO, improvement).
- Human effort to ship every feature and fix
- Token cost per PR, work item, and team.
Fig 1.1 — Shipped this week
Sample data
Bulk export
6 PRs · 2 follow-up bugs · 9.4 agent hr · 1.8 human hr · $212
Webhook delivery queue
3 PRs · no follow-ups · 2.1 agent hr · 0.6 human hr · $48
New billing module
4 PRs · 1 follow-up bug · 7.7 agent hr · 2.3 human hr · $174
- …and 22 other capabilities this week
| Work | New features | Bug fixes | Improvements |
|---|---|---|---|
| Merged PRs (48) | 42% · 20 PRs | 31% · 15 PRs | 27% · 13 PRs |
| Engineering hours (196 hr) | 42% · 82 hr | 20% · 38.5 hr | 39% · 75.5 hr |
| Agent spend ($1,240) | 42% · $520 | 17% · $210 | 41% · $510 |
| Bulk export | 70% | 20% | 10% |
| Webhook delivery queue | 10% | 85% | 5% |
| New billing module | 15% | 10% | 75% |
Fig 1.2 — Where the work went, 8 weeks
Sample data
- New features
- Improvements
- Bug fixes
| Week | New features (hr) | Improvements (hr) | Bug fixes (hr) | Bug fixes ($) |
|---|---|---|---|---|
| Week 1 | 14 | 16 | 22 | $290 |
| Week 2 | 16 | 16 | 20 | $265 |
| Week 3 | 17 | 16 | 19 | $240 |
| Week 4 | 18 | 17 | 17 | $215 |
| Week 5 | 19 | 17 | 15 | $190 |
| Week 6 | 20 | 18 | 14 | $165 |
| Week 7 | 21 | 18 | 12 | $145 |
| Week 8 | 22 | 20 | 10 | $130 |
How autonomous is the factory?
02 / Autonomy
In an effective software factory, the path from prompt to production is straight. Fewer mid-session corrections, less rework, fewer review cycles — and code that lands in production.
Critical metrics
- Session autonomy: the steering and interruptions required to get good code.
- PR rework required before merging.
- Amount of code discarded locally or in closed PRs.
- Differences in autonomy between local sessions, background agents, and automations.
Fig 2.1 — Prompt → production
Sample data
| Run | Human interventions | Steps |
|---|---|---|
| Low autonomy | 7 | clarify, steer, share docs, re-prompt, correct, review, rewrite |
| Medium autonomy | 4 | clarify, re-prompt, review, correct |
| One-shot | 1 | review |
Where does human time go?
03 / Human effort
In a healthy factory, agents give humans leverage — the ability to do more. So where is their time going? What are the new bottlenecks? And how are the most effective developers working?
Critical metrics
- Human effort and tokens for each PR.
- Waiting time: hours waiting on review, or on long-running agents.
- Time spent reworking PRs.
- Time spent reviewing PRs.
- Token spend per person, broken down by output.
Fig 3.1 — The human week
Sample data
| Hours in the loop | MON | TUE | WED | THU | FRI | Total |
|---|---|---|---|---|---|---|
| Waiting on agents or review | 2 | 2 | 2 | 2 | 2 | 10 |
| Reviewing PRs | 2 | 1 | 2 | 1 | 1 | 7 |
| Prompting and Steering | 3 | 3 | 3 | 4 | 2 | 15 |
| Reworking | 1 | 2 | 1 | 1 | 3 | 8 |
Fig 3.2 — AI maturity, per contributor
Sample data
| Contributor | Maturity (1–5) | Level |
|---|---|---|
| R. Okonkwo | 5 | Orchestrates |
| J. Lee | 5 | Orchestrates |
| M. Patel | 4 | Agentic |
| S. Rivera | 4 | Agentic |
| A. Chen | 4 | Agentic |
| T. Novak | 3 | Assisted edits |
| D. Haas | 3 | Assisted edits |
| K. Murphy | 2 | Q&A |
What does a unit of shipped work cost?
04 / Cost & ROI
Account for every human hour and every dollar your factory spends. Find what to optimize, name the real waste, and show a clear ROI for the work each team is doing.
Critical metrics
- Session yield by count and dollars: shipped to production, rework, discarded, closed PR.
- Cost per work type, with trends.
- Estimated agent hours from each model.
- Average cost per PR type.
- Dollars spent on code review.
Fig 4.1 — Session yield
Sample data
| Outcome | Sessions | Share | Cost |
|---|---|---|---|
| RESEARCH & BRAINSTORM | 114 | 9% | $1.3K |
| SHIPPED TO PRODUCTION | 512 | 43% | $18.4K |
| SHIPPED AFTER REWORK | 218 | 18% | $8.9K |
| DISCARDED | 214 | 18% | $2.9K |
| PR CLOSED UNMERGED | 146 | 12% | $1.2K |
Fig 4.2 — Leverage & unit cost
Sample data
- Est. human hours from AI
1,840 hr
+18%vs last quarter
- Avg. cost per feature
$41.11
−24%vs last quarter
- Avg. cost per bug fix
$16.11
−31%vs last quarter
- Avg. cost per code review
$21
−8%vs last quarter
| Measure | Current | Change vs last quarter |
|---|---|---|
| Est. human hours from AI | 1,840 hr | +18% |
| Avg. cost per feature | $41.11 | −24% |
| Avg. cost per bug fix | $16.11 | −31% |
| Avg. cost per code review | $21 | −8% |
Factory Maturity
Everyone is figuring out how AI fits their daily work. These are the stages we see individuals and teams move through.
Team
- 1Asking AI
Developers reach for AI to answer questions, but are not running agentic flows.
- 2AI writes the code
Developers use agent mode and most PR code is AI-written. The team largely operates the same as before.
- 3Central team tunes the agents
A central team drives the agentic SDLC — building custom skills, agents, governance, and MCPs to make agents more effective.
- 4Every team tunes the agents
Each team owns improving AI performance on their part of the codebase — the outer loop is everyone’s job.
- 5Prompt to production
There is a straight line from prompt to production. ~50% of work is kicked off from Slack or Linear, accept rates are high, and low-risk PRs automerge.
IC
- 1Tab-completer
Still working the way they did before — AI as autocomplete.
- 2Single agent
AI-aided development, working on one PR at a time.
- 3Multi-agent
Managing multiple agents and PRs at once, focusing on plans, tests, and specs to stay aligned with the agents.
- 4Harness engineer
AI writes every PR. When it’s not working, goes up a level and adjusts the harness.
Track AI Code all the way to production
Observability for your software factory.
Tracks code from every major coding agent
- Cursor
- Codex
- Claude Code
- GitHub Copilot
- Gemini
- OpenCode
- Junie
- Droid
- Windsurf
- Amp