Skip to content

Research · Benchmark · April 30, 2026

The Defensibility Rubric, 2026

The public scoring rubric behind every Codritium benchmark and certification: six weighted dimensions, twelve task categories, inter-rater agreement, and a frozen reference quarter.

Q2 2026 reference · v2.6
Codritium Research

The point of a public rubric

A benchmark that can't be replicated is marketing. The rubric below is the one we use to score every challenge on the platform — and every report we publish. It is frozen for the reference quarter and versioned thereafter.

6

Weighted dimensions

Sum to 1.0

12

Task categories

Coverage map below

0.81

Inter-rater agreement

Krippendorff's α, Hard tier

Q2 2026

Frozen reference

Re-versioned each quarter

The five-stage pipeline

Every scored session passes through five stages. The first four are automated; the fifth is panel-reviewable for Hard-tier sessions.

01 · Session
Candidate opens a challenge in the IDE.
02 · Solve
AI pair available. Edits, runs, iterates.
03 · Defend
Replay annotated. Decisions explained.
04 · Score
Six rubric dimensions, weighted.
05 · Verdict
Panel-reviewable for Hard tier.
The scored-session pipeline.Codritium Scoring v2.6

The six rubric dimensions

Weights sum to 1.0. Three correctness-shaped dimensions (left), three judgment-shaped dimensions (right). The shape is the rubric's most-debated property — earlier versions weighted correctness alone.

Rubric weights, normalized to 100.Codritium Scoring v2.6 vs legacy

Twelve task categories

DomainCategoryTasksTop tier
DebuggingSingle-service142Hard
DebuggingDistributed86Hard
SecurityAuth / authz71Hard
SecurityInjection / SSRF64Hard
RefactoringWithin-module118Medium
RefactoringCross-module92Hard
Feature buildBounded104Medium
Feature buildCross-cutting58Hard
System designWrite-path39Hard
System designRead-path41Hard
Code reviewAccept / reject86Medium
Incident responseDiagnosis47Hard
Full category map.Frozen for Q2 2026

Inter-rater agreement, before and after

Replay-grounded scoring nearly doubled agreement on Hard-tier panel reviews. Reviewers were calibrated on a 40-task warm-up before scoring counted toward the published number.

Krippendorff's α across rubric versions, Hard tier.Same 240 calibration tasks scored each version

Replicate this

The rubric source, calibration set, and rater training notes are versioned in the open. The next version (v2.7) freezes on the first business day of Q4 2026.