Why this matters
Most platform reports look at a snapshot. They tell you how a cohort scored last quarter, or how a feature shifted a metric in a window. They are useful and they are short-sighted. The harder question is what happens when you watch the same engineers for nine months — where they grow, where they stall, where they leave.
We had the rare opportunity to run that study. 318 engineers committed to a nine-month tracking window starting 2025-09-15. They consented to longitudinal anonymized analysis. They did not get coaching that the rest of the platform population didn't get. They just kept practicing, and we kept the lights on the dashboard.
What the data shows is more uneven than the marketing-friendly story. Some engineers improved on every rubric dimension. Some plateaued at Medium and never crossed into Hard. Some dropped out for reasons that had nothing to do with the platform. We tell that whole story here.
Methodology in one paragraph
A "tracked engineer" is one of the 318 who completed at least one scored session in each of the nine 4-week windows from 2025-09-15 to 2026-06-13. Engineers who missed any window for more than 21 consecutive days are categorized as "dropouts" and broken out separately. Skill-band assignment is made fresh per session by the platform's adaptive scheduler based on rolling rubric composite; it is not sticky. Subscore weights changed in 2026-02-14 when we adopted rubric v2.6; we re-scored all sessions in the tracking window retroactively under v2.6 so the trajectories are consistent. Comparisons across quarters use the v2.6 composite for all three quarters.
318
Engineers tracked
Consented, nine-month window
14,520
Sessions scored
Mean 45.7 per engineer
62%
Moved up a band
Easy → Medium or Medium → Hard
18.6%
Dropped out
Across all reasons
The skill-band ladder
The first picture is the cohort's session mix over time, by band. Weekly resolution, stacked area. It is the simplest answer to "did the cohort grow."
Three things stand out. Easy-tier sessions drop from 62% to 17% — the cohort outgrows the band where they started. Hard-tier sessions climb from 7% to 43%, but the climb is fastest in weeks 11–22, then slows. Medium holds remarkably steady at around 40% from week 13 onward. The cohort doesn't graduate from Medium so much as it adds Hard on top of it.
The slope chart of composite rubric score per band aggregate confirms the same thing. Top-band engineers gained the most in absolute terms; mid-band engineers gained the most in percentage terms. Bottom-band engineers gained the least, which is the inverse of what coaching-focused interventions usually produce — and it is a clue about where the plateau will appear.
The slope at the top (+0.71) is more than 2.5× the slope at the bottom (+0.27). That is not because the bottom decile is malicious or under-engaged. Their session counts are similar to the median. They are running into something else.
Where the gains came from
A composite score moving up is only useful if you can see which subscores moved. The heatmap below is per-quarter mean subscore for the cohort. We watch movement within a row, not across rows.
| Q1 | Q2 | Q3 | |
|---|---|---|---|
| Correctness | 3.41 | 3.62 | 3.74 |
| Regression cost | 2.89 | 3.18 | 3.31 |
| Defensibility | 2.54 | 3.04 | 3.58 |
| Spec fidelity | 2.71 | 2.88 | 3.01 |
| AI-rejection | 2.38 | 2.81 | 3.34 |
| Replay clarity | 2.81 | 3.21 | 3.62 |
Correctness moves first and least. The cohort entered already at 3.41 out of 5 on correctness; they exited at 3.74. A 0.33-point move is real but small. Defensibility moves last and most — 2.54 → 3.58, more than a full rubric point. The pattern is the one we keep seeing across the platform: correctness comes early because the rubric rewards working code, and working code is what most engineers practice first. Defensibility comes later because it requires the engineer to look back at their own reasoning, which is a skill people have to be taught they are allowed to develop.
The multi-series line chart over 36 weeks shows the timing more crisply. The Defensibility curve is below every other curve until week 14 and then climbs through week 32 to become the second-highest curve.
Three subscores plateau early: Correctness, Regression cost, and Spec fidelity. The last of these is the one we want to draw attention to. Spec fidelity moves 0.30 across the whole nine months. Engineers who never learn to push back on under-specified prompts stay under-specified for a long time. The plateau in Spec fidelity is the plateau most engineers in the cohort hit and don't get out of.
The plateau pattern
A reasonable question at this point is what stops the median engineer at Medium. The composite score data and the subscore data converge on the same answer. Most engineers in the cohort hit a point — usually around week 18 to week 24 — where they have learned to ship correct code under the rubric and have not yet learned to defend it. They cycle through Medium sessions, score around 3.0 on correctness, and gain almost nothing on Spec fidelity or Defensibility for a stretch of weeks.
The engineers who break out of the plateau do one of two things. They run through the structured replay exercises (introduced in week 13 for the cohort) and start using replay to interrogate their own decisions. Or they get matched into a Hard-tier panel-reviewed session and the panel's feedback on Defensibility produces a sharp lesson. Either path is enough; doing both is best. The engineers who don't do either keep cycling at Medium until they either commit to defensibility-shaped practice or stop coming.
This is not a story about willpower. It is a story about which mechanism makes the rubric's later dimensions legible. The data says replay is the mechanism that works.
The dropout anatomy
18.6% of the cohort dropped out before the nine-month window closed. That number alone is not very useful. The reasons matter. We coded the 59 dropouts by exit-survey response (52 responded) plus inferred reason from session activity (the remaining 7).
The largest single reason — 21 engineers — is "took an offer and paused practice." That is a good outcome for the engineer and a confounder for the analysis; we don't get to keep watching them grow once their loop closed. The second-largest, 14 engineers, is "plateaued and disengaged." Those are the ones who hit the Medium-band wall and lost momentum. The remaining categories are mostly orthogonal to platform performance — life happened, careers changed, or an employer policy disallowed external practice tools.
The dropout pattern matters for the cohort statistics in the rest of the piece. The 318 remaining at week 36 are not the original 318. They are the 259 who didn't drop out. Survivor bias is real here. The composite-score climbs we report would be smaller — perhaps 10–15% smaller in slope — if we backfilled the dropouts with their last-known composite carried forward.
The standout improvers
Twelve engineers in the cohort gained more than 1.5 composite points across the three quarters. That is more than 3× the cohort median gain (0.42). We pulled their starting and ending bands and the sparkline of their weekly composite. Anonymized; the rank is by Δ composite.
| # | Engineer | Start band | End band | Δ composite | Weekly trajectory |
|---|---|---|---|---|---|
| 1 | E-104 | Easy | Hard | +2.18 | |
| 2 | E-217 | Easy | Hard | +2.04 | |
| 3 | E-082 | Medium | Hard | +1.94 | |
| 4 | E-138 | Easy | Medium | +1.81 | |
| 5 | E-046 | Medium | Hard | +1.76 | |
| 6 | E-271 | Easy | Medium | +1.72 | |
| 7 | E-159 | Medium | Hard | +1.68 | |
| 8 | E-193 | Easy | Hard | +1.64 | |
| 9 | E-302 | Medium | Hard | +1.62 | |
| 10 | E-029 | Easy | Medium | +1.58 | |
| 11 | E-251 | Medium | Hard | +1.54 | |
| 12 | E-067 | Easy | Medium | +1.51 |
The sparklines look similar by design. The standout improvers do not zig-zag. They climb steadily. We pulled the session-level data for these twelve and found three behaviours that distinguish them from the cohort median: their replay-clarity score moves above 3.0 by week 10 (median: week 21), they spend disproportionately more time on Defensibility-rich categories like security and refactoring (median ratio 1.8× the cohort), and their AI-rejection rate is consistently above 0.4 (median engineer reaches that threshold by week 24, the standouts by week 8).
We are not claiming the path generalizes. We are saying: the engineers who improved the most started practicing replay early and never stopped.
Milestones, when the median engineer hits them
A few benchmark events tell the story of the median engineer's nine months.
The gap between week 11 (first Hard attempt) and week 18 (first Hard pass) is the plateau period. Seven weeks of attempting and not yet clearing. The cohort that gets the most out of this stretch is the cohort that uses it for replay-grounded reflection. The cohort that treats it as a series of failed attempts often drops out somewhere inside weeks 13–22, which is exactly when the "plateaued and disengaged" cluster of the dropout donut concentrates.
Caveats
This is not randomized. The 318 self-selected into a nine-month tracking commitment, which means they were already more engaged than the platform median when the window opened. The composite-score climbs are real but probably overstate what a random engineer would gain over nine months.
Survivor bias on the 259 remaining at week 36 is real. We report cohort statistics on the surviving group, not the original 318. We checked whether backfilling dropouts with last-known-composite carried forward changed the conclusions; it shrinks the absolute gains by ~12% but does not change the rank-order of subscore movement, so we kept the survivor-only presentation as cleaner.
Rubric v2.6 launched mid-tracking (2026-02-14). We re-scored all sessions retroactively under v2.6 so the trajectory is consistent. Engineers who got an earlier session re-scored may have seen the score move in a panel-confirmed way; we did not back-propagate panel-review verdicts because the v2.6 rubric also moved which sessions are panel-eligible. The methodology appendix has the full audit.
The plateau-pattern interpretation is supported by the data but the mechanism (replay practice unlocks Defensibility) is correlational. We are launching a randomized A/B in H2 2026 — half the cohort gets a structured-replay nudge at week 13, half doesn't — to test the mechanism causally.
What we'll do next
We re-run this study with the next nine-month cohort starting 2026-09-15. The new cohort gets the replay nudge experiment described above. We also add a per-engineer skill-band confidence interval (currently bands are point-assigned per session; the new version gives a 95% band over the rolling 12 sessions). The data dictionary, anonymized weekly composites, and v2.6 rubric snapshot are versioned in the methods repo. Independent groups running similar cohort studies — particularly hiring teams running 90-day onboarding programs — are welcome to compare. The full rubric is in the Defensibility Rubric, and the cross-quarter regression-rate cuts are in Regression Rate by Skill Band.