Skip to content

Blog · June 7, 2026 · 8 min read · by Codritium Editorial

2026 in Engineering: The Year AI Hit the On-Call Rotation

A narrative recap of 2026 in software engineering — the year AI moved from autocomplete to on-call rotation, and the regression-rate data caught up with the rhetoric.

recap
year-in-review
outlook

In 2024 the question was whether AI could write the code. In 2025 the question was whether the code was correct. In 2026 the question changed shape. AI moved out of the editor and into the on-call rotation — drafting runbook entries, triaging alerts, proposing rollbacks — and the engineering culture had to decide what kind of teammate it actually was. The decision is not finished. The data, however, has started to settle. This is the year as we read it, told through the events, the numbers, and the lines we keep hearing back from the teams we work with.

The year in five beats

Frontier model release
Feb 2026
Long-context coding model from a major lab; pushes paired-session productivity ceiling up about 18%.
Regression-rate paper
Mar 2026
Public release of the six-quarter panel showing top-tercile regression dropped to 2.8%.
Rubric v2.6
May 2026
Defensibility anchors tightened; AI-rejection sense raised to 12%. Inter-rater agreement up to 0.81.
On-call pilots
Aug 2026
Three FAANG-class employers begin AI triage pilots for Sev-3 and Sev-4 alerts.
Panel-eligible listing
Oct 2026
First public senior job listing requesting 'panel-eligible' candidates. Earlier than we forecast.
On-call benchmark
Dec 2026
Codritium publishes the first benchmark of AI vs human triage accuracy on a fixed alert set.
The year's major events as we read them. Six anchors that shaped the rest of the field.Codritium Research, 2026 year-in-review.

What changed about how we ship code?

The shape of a normal commit changed in 2026. Two years ago, AI suggested lines and an engineer accepted them inside their editor; the diff still looked like a person typed it. This year, the dominant pattern is the AI proposes a multi-file change and the engineer reviews it before merging. The review is the work. The acceptance criterion is no longer "does it compile" but "do I understand why this is the change to make." That shift is small in words and large in practice. It moved engineering work toward review and away from authorship, and it produced a measurable redistribution of where engineers spend their hours.

The on-call rotation followed the same arc roughly six months later. By Q3 2026, the dominant pattern for Sev-3 and Sev-4 alerts at large employers was: AI triages the alert, drafts the rollback or the runbook entry, and routes the result to a human on-call who approves or rejects within a defined window. The pattern is not universal — it covers something like a third of organisations we surveyed — but it is the leading edge. Sev-1 and Sev-2 stayed human-driven; nobody wants an AI to be the last approver on a database failover, and that taste appears to be holding.

The thing nobody warned us about is that the AI is fast at draft and slow at judgment. Triage at 3am is mostly judgment. So we use it as a first reader, not a first responder. That distinction matters.

Staff engineer, payments platform

What changed about how we interview?

The interview loop in 2026 looked different at the senior level than at any point in the previous five years. The defence audio — a five-to-ten minute asynchronous explanation of decision points, paired with the take-home — went from rare to common in eighteen months. Forty-seven percent of FAANG-class senior screens included a defence audio requirement by Q2; that number was effectively zero in early 2024. The mechanism was not subtle. Employers wanted a way to tell the artifact apart from the engineer, and audio was the cheapest available answer.

The rubric refresh in May was smaller than we expected. We had predicted AI-rejection sense would jump from 10% to 14% of the defensibility score; it moved to 12%. The panel reviewers were more conservative than the data alone justified, and the conversation that produced that 12% was about calibration stability rather than predictive power. We think 22% by late 2027 is still the right trajectory; the May refresh was a pause, not a reversal.

The October job-listing surprise was real. We had forecast the first public 'panel-eligible' senior listing would arrive in Q1 2027. A mid-sized infrastructure employer posted one in October 2026 — a full quarter early. Two more followed by year-end. The early adoption shifted our internal forecasts for 2027 employer adoption upward by roughly half a quarter.

What changed about how we measure?

The single biggest measurement shift of 2026 was the regression-rate paper in March. The number itself — top-tercile regression rate falling from 4.2% to 2.8% over six quarters — got more citations across industry blogs and engineering newsletters than any other engineering statistic we tracked. The reason is that it answered a question senior engineers had been asking without a clean answer: are the tools actually helping the strong engineers, or are they pulling everyone toward the same mediocre middle? The data said the tools amplified the strong engineers and barely touched the weak ones, and that asymmetry became the central fact about hiring this year.

The AI-rejection rate followed roughly a quarter later. Until 2026 it had lived as a rubric input — a thing reviewers scored, but not a number anyone reported on its own. The Q3 release of cohort-level AI-rejection distributions changed that. Engineers started quoting their own rejection rate the way they had previously quoted their commit cadence, and the number became part of the working vocabulary of senior practice. The risk of any metric becoming public is that it gets gamed; we saw some of that. The countermeasure was the audio defence pairing — gamed rejection rates produce inconsistent defences, and the panels catch the inconsistency.

The on-call benchmark in December was the year's quieter measurement story. Codritium published the first apples-to-apples benchmark of AI triage accuracy against human triage on a fixed alert set. The headline finding — AI matched senior on-call accuracy on Sev-3 and Sev-4 alerts and underperformed by a wide margin on Sev-1 and Sev-2 — gave the field a number for an intuition that had been growing all year. Nobody was surprised by the direction. The magnitude of the Sev-1 gap was larger than most teams had assumed, and it shaped how Q1 2027 budgets got written.

What didn't change (that we expected to)?

We expected three things to move in 2026 that did not. The take-home assignment did not rebound at the senior level the way we forecast; it sat at 32% adoption, roughly flat with 2025. The rebound is still coming, but it is coming through 2027, paired with the defence audio. Our 2026 forecast was a year early. The mechanism is right; the calendar was wrong.

We held off on take-homes through 2026 because the defence audio tooling was not ready until late summer. We are running them now. The reason we waited was nobody wanted to be the first employer holding a take-home that turned out to be the candidate's tool, not the candidate.

Hiring lead, infrastructure company

The security-review workflow stayed mostly unchanged, which surprised us. We had expected the SSRF and prompt-injection findings from the late-2025 research wave to produce a structural change in how teams review AI-touching code. They produced a stylistic change — more checklists, more linters — and not a structural one. The mechanism we underweighted was inertia. Security review changes follow incidents, not papers, and 2026 did not produce a headline incident that forced the structural change.

And the junior engineer pipeline did not pivot. We thought bootcamps and university programs would update their curricula faster than they did. They didn't. The structural pivot is happening at the company-onboarding level, not at the schooling level. Whether that holds in 2027 is one of the open questions.

What we got wrong in our own predictions

Three of our late-2025 forecasts for 2026 missed. The panel-eligible listing arrived in October instead of Q1 2027 — we were wrong by a quarter in the right direction, which is the most flattering kind of miss. The take-home rebound did not happen at all in 2026; we were a full year early. And the rubric refresh in May moved AI-rejection sense to 12% rather than the 14% we had forecast — a smaller magnitude than the data alone justified, because calibration stability won the room.

The honest assessment is that two of those misses are about timing and one is about magnitude. We are better at direction than calendar. That asymmetry is consistent with our 2025 retrospective, where directional hits outpaced magnitude hits by roughly two to one. The mechanism is that we read the data well and read the institutional adoption cycles less well. We will be slightly more conservative on adoption-cycle timing in the 2027 forecast set, and we will publish the gap between our internal "data says" forecast and our "we expect the institutions to" forecast separately. Two numbers each, not one.

If 2024 was the year we asked whether AI could write the code, and 2025 was the year we asked whether the code was correct, 2026 was the year we asked who was responsible for it. We are still answering.

Codritium Research, December 2026

Common questions

  • Is this a US-only view?

    Largely, weighted toward FAANG-class employers in the US and a handful of UK and EU mid-sized companies. The on-call adoption number is roughly US-only. The regression-rate panel includes engineers from twelve countries but skews North American. We do not have enough APAC representation to claim global coverage.

  • What about open source?

    Open-source maintenance patterns shifted less than corporate engineering in 2026. AI is used heavily for first-pass triage on issues, less for accepting PRs into the trunk. The maintainer's review-as-work pattern existed long before AI made it mainstream in corporate engineering; open source led the curve by a decade and did not need to pivot the same way.

  • Did employer adoption hit your forecast?

    Partially. The defence audio adoption hit our forecast within a percentage point. The panel-eligible listing came a quarter earlier than we forecast. The take-home rebound did not arrive in 2026 at all; it slipped into 2027 conditional on defence-audio pairing. Our forecast track record is better on direction than on calendar.

  • Are these your own data or industry-wide?

    A mix. The regression-rate panel and the AI-rejection distributions are ours, drawn from our cohort and disclosed in the research collection. The on-call adoption numbers are from a survey we ran with 1,840 respondents across 12 countries. The job-listing data is from a tracker we maintain that scrapes public listings from a fixed list of employers. We try to disclose the source on each number; let us know if any are unclear.

  • What was 2026's biggest miss?

    Our forecast for the take-home rebound. We had it at 38% by Q4 2026; actual was 32%, basically flat. The mechanism we underweighted was that the defence-audio tooling that makes paired take-homes honest was not production-ready until late summer. Employers held off rather than running take-homes without the pairing. The miss is about reading institutional caution, not the data.

Where to go next