Predictions are accountable. We are publishing these with the dates, the numbers, and the conditions we expect them to land under, so we can be told in December 2027 which ones missed. The five below are the ones we are most willing to defend. None of them are safe. Two of them are deliberately uncomfortable. The confidence section at the end is honest about which ones we would already revise.
Why predictions, not just data
We publish a lot of data. Data describes what already happened. Predictions describe what we expect to happen and the mechanism by which we expect it. That mechanism is the part we can check. A wrong number with a right mechanism teaches us something; a right number with a wrong mechanism teaches us nothing and quietly makes us worse forecasters next year. We are publishing these to be told which mechanisms are wrong. The rest is bookkeeping.
Prediction 1: The rubric's AI-rejection rate weight will rise from 12% to 22% by Q4 2027
The current rubric assigns 12% of the defensibility score to AI-rejection sense — the engineer's demonstrated ability to refuse a plausible AI suggestion for a reason. The weight was conservative when it was set in early 2024, because we did not yet have a year of regression-rate data tied to rejection behaviour. We do now. The correlation between rejection rate and downstream regression-rate is stronger than any other single rubric input, and it has held across two rubric refreshes. The natural next move is to weight the input proportional to its predictive power.
The path to 22% is incremental. We expect the v2.7 refresh in late 2026 to move it to 16% and the v2.8 refresh in mid-2027 to move it to 22%. Both increments will be small enough that calibration agreement does not break; together they nearly double the weight. The mechanism is simple — measure what predicts, weight it accordingly — and the cost is borne by engineers who optimise for clean diffs without demonstrating refusal behaviour. That cost is the point.
Prediction 2: At least one FAANG-class loop will publicly require an AI-rejection log
The current interview loop at most FAANG-class employers measures what the candidate produced. Half a dozen employers we know of are quietly piloting a second artifact: a log of suggestions the candidate considered and refused, with a one-line reason per refusal. The pilots are private because no employer wants to be first publicly. We think one will be public by Q3 2027. The mechanism is competitive — once one employer ships a rejection log requirement and links it to retention data, the others will follow within two cycles to avoid an adverse selection on rejection-sensitive candidates.
The candidate behaviour this will create is mixed. Some candidates will over-log, listing rejections that did not happen, which the panels will catch through audio defence inconsistency. Others will under-log, which the panels will catch by asking about a specific decision point with no rejection note. The selection pressure is, however, real. Within a year of the first public adoption, "I track my rejections" becomes a normal answer to "how do you work with AI."
| Quarter | Milestone | Confidence |
|---|---|---|
| Q4 2026 | Three or more private pilots running at FAANG-class employers | High |
| Q2 2027 | First public mention of rejection-log requirement in a job posting | Medium |
| Q3 2027 | At least one employer publicly mandates the artifact | Medium |
| Q1 2028 | Three or more employers require the artifact for senior loops | Low |
Prediction 3: The regression-rate gap between top and bottom terciles will widen
The popular view is that better tools narrow the gap between strong and weak engineers because the floor rises faster than the ceiling. The data so far does not support that view. Across our six-quarter regression-rate panel, the top tercile's regression rate fell from 4.2% to 2.8%. The bottom tercile's fell from 11.8% to 11.1%. The gap widened from 7.6 percentage points to 8.3, not because the bottom got worse, but because the top got better faster. Better tools amplified the engineers who already had the habits to refuse, verify, and reproduce. They did not transfer those habits to engineers who didn't.
We expect this divergence to continue through 2027. By Q4 2027 we expect the top-tercile regression rate to fall to about 2.1% and the bottom-tercile to sit at about 10.8%. The gap is the headline number. It is also the most important fact about hiring in 2027 — the cost of hiring from the bottom tercile, in regression dollars, will be more than three times the cost of hiring from the top.
Prediction 4: 'Panel-eligible' will become a hireable signal
The shape of this prediction is the most speculative of the five. The mechanism is signalling: as the cost of evaluating senior engineers rises, employers will reach for any pre-vetted signal that reduces that cost. A passed panel review is a strong such signal because the rubric is public, the inter-rater agreement is measured, and the artifact is portable. We expect the first job postings asking for "panel-eligible" or equivalent language to appear in Q1 2027 and to be common in senior listings by Q4 2027. Common means "appears in roughly thirty percent of senior listings at FAANG-class employers and at the better-organised mid-sized companies."
The risk in this prediction is not the direction; it is the timing. Signalling adoption tends to be slower than we forecast because employers are reluctant to add hiring requirements that narrow the pool. We may be a year early. The mechanism is still the right one — the pressure to use pre-vetted signals will only rise — but the calendar may slip into 2028 for the "common" milestone.
Prediction 5: The take-home assignment will rebound
Take-homes were the dominant senior-loop format until 2024, then collapsed because employers worried they were measuring the candidate's AI tools more than the candidate. The collapse was real — take-home use in senior loops fell from about 62% in 2023 to 28% in 2025. The rebound is already starting because the alternative formats have not held up. Live coding measures speed, not judgment. Whiteboard system design measures rehearsal. Panel reviews measure judgment but cost more per candidate than employers want to spend at the screen stage.
The rebound is conditional. Take-homes return paired with a defence audio — five to ten minutes of the candidate explaining their decision points, recorded asynchronously. That pairing puts the artifact under reasoning audit, which is exactly the gap that killed the unpaired take-home. We expect take-home + audio adoption to reach roughly 45% of senior loops by Q4 2027.
How confident are we?
The confidence chart below shows where we would actually put money. The rubric prediction is high-confidence because we control the rubric and the mechanism is already running. The regression-rate divergence is high-confidence because the data has been moving in one direction for six quarters with no inflection. The take-home rebound is medium-high because the alternative formats are not holding up. The AI-rejection log is medium because the pilots are real and the adoption pressure is real, but the announce date is hard to forecast. The panel-eligible signal is the one we would already revise — likely right, plausibly a year late.
If only one of these lands, we hope it is the regression-rate divergence. The cost of getting that one wrong is paid by the engineers who got hired into roles their tools could not actually carry them through.
The longitudinal cohort study with the regression-rate divergence data underlying Prediction 3.
From Codritium Research
Three-Quarter Cohort Drift
Tracking one cohort across three quarters. Where 318 engineers improved, where they plateaued, where they regressed. Skill-band ladders, dropout patterns, and the rubric subscores that moved most.
Common questions
Why publish predictions instead of just data?
Because data describes the past and predictions check the model behind it. A wrong number with a right mechanism teaches us something; a right number with a wrong mechanism teaches us nothing and quietly degrades our forecasting next year. We publish to be told which mechanisms are wrong.
What's the track record?
Our 2025 predictions hit four of six on direction and two of six on magnitude within their stated time windows. The two magnitude misses were on regression rate (we underestimated the divergence) and on competition adoption (we overestimated). The full 2025 retrospective is in the State of AI Engineering 2026 report.
Are these company-level or industry-level?
Industry-level for predictions 1, 3, and 5. Employer-level for predictions 2 and 4. We have higher confidence on the industry-level forecasts because they are slow-moving and underpinned by data we collect directly. Employer-level forecasts are inherently noisier because they depend on individual adoption decisions.
How will you measure success?
Each prediction has a specific number and a specific date. We will publish a retrospective in January 2028 grading each as hit on magnitude, hit on direction only, or missed. We will not move the goalposts. If a prediction misses by 18 months, that is a miss, not a deferred hit.
What's the one prediction you'd revise today?
Prediction 4 — the panel-eligible signal. We believe the direction is right but the 2027 timing is aggressive. We would now put roughly 60% odds on the milestone landing in 2028 instead. We are leaving the published prediction as written because moving the goalposts on day one defeats the purpose.
Where to go next
- The longitudinal data behind Prediction 3: Three-quarter cohort drift
- The current rubric Prediction 1 forecasts against: The defensibility rubric, 2026
- The format Prediction 5 forecasts the rebound of: Take-home assignments 2026: standout guide