Stack Evals

Which model earns which job in the harness. Every big action gets its own lane of evals — triage, review, build, judgement — at graded complexity (T1 mechanical → T3 subtle), with the house engineering rules present in each workspace and the production prompts where the harness has them. Hidden assertions grade the behavior, not the promise; the routing table is the output that feeds each OS's model config.

Updated 2026-08-21 · fork of vercel/next-evals-oss
9models scored
157/170house cells passed
19house evals across 4 lanes
24upstream Next.js evals

Triage lane

Reads the ticket, the thread, and the board the way a senior engineer does before anyone writes code — including the traps: unanswered questions, mid-thread reversals, near-duplicate bait. Runs the production triage prompt.

Model Score Answered → readyT1 Unanswered trapT2 Mid-thread reversalT3 Cost/eval
Gemini 3.1 Pro PreviewGemini CLI 3/3 $0.11
Grok 4.6OpenCode 3/3 $0.15
Claude Sonnet 5Claude Code 3/3 $0.34
GPT 5.6 Sol (ultra)Codex 3/3 $0.71
MiniMax M3OpenCode 2/3 $0.03
Kimi K3OpenCode 2/3 $0.13
GLM 5.2OpenCode 2/3 $0.23
Claude Fable 5 (high)Claude Code 2/3 $0.94
Claude Opus 5Claude Code 2/3 $1.22

Review lane

Pre-merge call on a PR diff with seeded defects from real blocked builds (cited by D-rule). Graded on recall AND precision — a clean diff must be approved.

Model Score Clean diffT1 Metered defectsT2 Hidden violationsT3 Cost/eval
MiniMax M3OpenCode 3/3 $0.03
Gemini 3.1 Pro PreviewGemini CLI 3/3 $0.11
GLM 5.2OpenCode 3/3 $0.23
Claude Sonnet 5Claude Code 3/3 $0.34
Claude Fable 5 (high)Claude Code 3/3 $0.94
Kimi K3OpenCode 2/3 $0.13
Grok 4.6OpenCode 2/3 $0.15
Claude Opus 5Claude Code 2/3 $1.22
GPT 5.6 Sol (ultra)Codex 1/3 $0.71

Build lane

PR-shaped coding tickets at graded sizes, T1 mechanical through T3 multi-law feature, with the house rules in the workspace. Hidden vitest assertions grade the shipped behavior.

Model Score Confirm modalT1 RLS migrationT1 Fix red CIT1 house-062-pr-pagination-commitT1 Metered dedupT2 Browse pageT2 DB-driven modelT2 house-063-pr-jitter-digestT2 Usage ledgerT3 house-064-pr-funnel-healthT3 Cost/eval
Kimi K3OpenCode 10/10 $0.13
Grok 4.6OpenCode 10/10 $0.15
GLM 5.2OpenCode 9/9 $0.23
Claude Sonnet 5Claude Code 10/10 $0.34
Claude Fable 5 (high)Claude Code 10/10 $0.94
Claude Opus 5Claude Code 10/10 $1.22
MiniMax M3OpenCode 9/10 $0.03
Gemini 3.1 Pro PreviewGemini CLI 9/10 $0.11
GPT 5.6 Sol (ultra)Codex 9/10 $0.71

Judgement lane

The accept/bounce call on finished work: is a zero-work success proof? Is green-but-unexercised done? Which approach survives the doctrine?

Model Score Dead-feed smellT1 Green ≠ doneT2 Approach choiceT3 Cost/eval
MiniMax M3OpenCode 3/3 $0.03
Gemini 3.1 Pro PreviewGemini CLI 3/3 $0.11
Kimi K3OpenCode 3/3 $0.13
Grok 4.6OpenCode 3/3 $0.15
GLM 5.2OpenCode 3/3 $0.23
Claude Sonnet 5Claude Code 3/3 $0.34
GPT 5.6 Sol (ultra)Codex 3/3 $0.71
Claude Fable 5 (high)Claude Code 3/3 $0.94
Claude Opus 5Claude Code 3/3 $1.22

Routing table

Who should run what: per lane and complexity tier, the cheapest model with a perfect cell score, and the runner-up. This is the table that feeds each OS's AgentConfig — a model change is a data edit, never a deploy. Cells with small n are indicative, not decisive.

LaneTierRecommendedRunner-upn
triage T1 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
triage T2 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
triage T3 Gemini 3.1 Pro Preview $0.11 Grok 4.6 $0.15 1
review T1 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
review T2 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
review T3 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
build T1 Gemini 3.1 Pro Preview $0.11 Kimi K3 $0.13 4
build T2 MiniMax M3 $0.03 Kimi K3 $0.13 4
build T3 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 2
judgement T1 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
judgement T2 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1
judgement T3 MiniMax M3 $0.03 Gemini 3.1 Pro Preview $0.11 1

What broke (first board, 2026-08-20)

gemini-3.1-pro · db-driven-model · 4/4 runs

Hardcoded the model ID anyway

Built the runtime config resolver correctly, then left a hardcoded claude-* literal at the feature call site — every attempt.

minimax-m3 · rls-migration · 1/2 runs

Forgot row-level security

Created the invoices table without ENABLE ROW LEVEL SECURITY on one attempt.

gpt-5.6-sol · browse-page · 1/2 runs

Table without its scroll container

Sortable table with linked rows, but no overflow container on one attempt.

Upstream Next.js board

The suite this repo forks from — framework-competency tasks. Includes cached results published by upstream, refreshed as evals change.

Model Score 000 app router migrati 021 avoid fetch in eff 022 prefer server acti 023 avoid getserversid 024 avoid redundant us 025 prefer next link 026 no serial await 027 prefer next image 028 prefer next font 029 use cache directiv 030 app router migrati 031 proxy middleware 032 use cache directiv 033 forbidden auth 034 async cookies 035 connection dynamic 036 after response 037 updatetag cache 038 refresh settings 039 indirect proxy 040 instant 041 optimize ppr shell 042 enable ppr 043 view transitions Cost/eval
Cursor Composer 2.5Cursor 23/23 $0.05
Kimi K2.6OpenCode 23/23 $0.06
Kimi K2.7 CodeOpenCode 23/23 $0.06
GLM 5.1OpenCode 23/23 $0.08
Claude Sonnet 4.6Claude Code 23/23 $0.14
Claude Opus 4.6Claude Code 23/23 $0.20
Claude Opus 4.8Claude Code 23/23 $0.35
Claude Opus 4.7 (max)Claude Code 23/23 $0.45
MiniMax M3OpenCode 23/24 $0.03
Grok 4.5OpenCode 23/24 $0.07
Kimi K3OpenCode 23/24 $0.13
GLM 5.2OpenCode 23/24 $0.23
Claude Fable 5 (high)Claude Code 23/24 $0.94
Cursor Composer 2.0Cursor 22/23 $0.04
GPT 5.3 Codex (xhigh)Codex 22/23 $0.23
Gemini 3.1 Pro PreviewGemini CLI 22/24 $0.11
Grok 4.6OpenCode 22/24 $0.15
Claude Sonnet 5Claude Code 22/24 $0.34
GPT 5.6 Sol (ultra)Codex 22/24 $0.71
Claude Opus 5Claude Code 22/24 $1.22
GPT 5.4 (xhigh)Codex 21/23 $0.24
Cursor Composer 1.5Cursor 21/23
Claude Sonnet 4.5Claude Code 20/23 $0.13
Gemini 3.0 Pro PreviewGemini CLI 20/23 $0.17
GPT 5.5 ProCodex 20/23 $18.21
GPT 5.2 Codex (xhigh)Codex 19/23 $0.15
MiniMax M2.7OpenCode 15/23 $0.03
Kimi K2.5OpenCode 14/23 $0.02