The point of a public rubric
A benchmark that can't be replicated is marketing. The rubric below is the one we use to score every challenge on the platform — and every report we publish. It is frozen for the reference quarter and versioned thereafter.
6
Weighted dimensions
Sum to 1.0
12
Task categories
Coverage map below
0.81
Inter-rater agreement
Krippendorff's α, Hard tier
Q2 2026
Frozen reference
Re-versioned each quarter
The five-stage pipeline
Every scored session passes through five stages. The first four are automated; the fifth is panel-reviewable for Hard-tier sessions.
The six rubric dimensions
Weights sum to 1.0. Three correctness-shaped dimensions (left), three judgment-shaped dimensions (right). The shape is the rubric's most-debated property — earlier versions weighted correctness alone.
Twelve task categories
| Domain | Category | Tasks | Top tier |
|---|---|---|---|
| Debugging | Single-service | 142 | Hard |
| Debugging | Distributed | 86 | Hard |
| Security | Auth / authz | 71 | Hard |
| Security | Injection / SSRF | 64 | Hard |
| Refactoring | Within-module | 118 | Medium |
| Refactoring | Cross-module | 92 | Hard |
| Feature build | Bounded | 104 | Medium |
| Feature build | Cross-cutting | 58 | Hard |
| System design | Write-path | 39 | Hard |
| System design | Read-path | 41 | Hard |
| Code review | Accept / reject | 86 | Medium |
| Incident response | Diagnosis | 47 | Hard |
Inter-rater agreement, before and after
Replay-grounded scoring nearly doubled agreement on Hard-tier panel reviews. Reviewers were calibrated on a 40-task warm-up before scoring counted toward the published number.
Replicate this
The rubric source, calibration set, and rater training notes are versioned in the open. The next version (v2.7) freezes on the first business day of Q4 2026.