Production case study

46 sprints of real velocity.

Measured Jira throughput from a real production application as Q moved from landing to scaling, compounding output and post-release maintenance.

From 7 developers to 1 engineer + Q
- output increased by 158%

A real production handover measured across 46 sprints.
One engineer running Q delivered higher output in a live production environment.

Original 7-person development team

Before Q, the company used a 7-person development team working on a two-week sprint cadence to maintain and advance the application.

Replaced by 1 engineer + Q

After a two-sprint transition, 1 engineer + Q took over development and moved to weekly releases. Within the first full month, the new model was already delivering at or above the prior team's multi-month output.

Over 46 sprints

We continued measuring the system through 46 production sprints, comparing the original 7-person model with 1 engineer + Q using real delivery data cross-referenced between Jira and GitHub.

91 → 205 weekly story points

Normalized to a 7-day sprint, average output increased from 91 to 205 story points - a 125% increase in weekly delivery efficiency.

643 → 1,658 monthly story points

Monthly completed work grew from 643 story points in April to 1,658 in July - a 158% increase while operating with 1 engineer + Q.

+19% • +28% • +69% MoM gains

The gains continued month over month: +19% from April to May, +28% from May to June, and +69% from June to July. Q didn't just replace team capacity - its delivery efficiency continued improving as the system matured.

Sprint-by-sprint throughput

Each bar represents completed story points for one measured sprint. Hover a bar for the point total.

What you're looking at ▾

Each bar is one sprint's completed story points. The chart contains 46 completed sprints, spanning S10 through S55 - about nine months of continuous delivery. The dashed line inside each phase is that phase's average; the white line is the monthly cumulative total, read against the right-hand axis.

The phases show the progression from manual development, to Q-led orchestration, to scaling, to compounding output through the release push, and finally into post-release maintenance. Sprint point totals are the team's own estimates, tracked in Jira and pulled directly from sprint reports - nothing here is modeled or projected.

The first phase (S10-S17) reflects a 7-person development team working the backlog by hand. Starting at LANDING (S18), that team was replaced by a single senior engineer operating Q. Through April, May, and June, the prompts and orchestration workflow were refined; by July and early August, those gains compounded into a step change in throughput and sprint cadence.

S43 marks the end of the major release push. From S44 onward, the backlog had largely shifted into maintenance mode. The cadence changed with it: instead of larger weekly sprints, work could be broken into multiple smaller sprint cycles within the same day or across a few days, with related items grouped into final pushes on similar dates. The post-S43 bars therefore represent a different operating mode - faster, smaller, and more targeted maintenance cycles rather than the same large development backlog seen earlier.

During the active development and release period: April → May +19%, May → June +28%, and June → July +69% in completed story points. August remained strong overall, but should be read as a post-release maintenance month, not as a direct continuation of the earlier growth curve.

Hover any bar for that sprint's exact point total and phase.

Phase average Monthly cumulative
Sprint points
Cumulative points
Sprint

What the benchmark proves

Not that story points are comparable across every organization - but that the same real application, team context and Jira estimation system changed materially as the operating model changed.

Production proof

Built on real work

This wasn't a coding benchmark or a controlled exercise. Q was used to deliver and maintain a live production application against the company's actual backlog.

Sustained results

Not a one-month spike

The results were measured across 46 production sprints, through active development, release, scaling, and maintenance. The gains continued as the operating model matured.

Repeatability

The second proof is underway

A second company is now using Q to build a new application while we expand the system around their Theme → Epic → Feature → User Story workflow - testing whether the gains can be reproduced in a different environment.