← CatalogueFrom Copilot to Colleague · QualityReader →

What the AI judges see

Every chapter is scored by the MASH judges across six dimensions — three of craft (humanness, voice, usefulness) and three of epistemics (evidence density, claim defensibility, non-redundancy). Higher is better; colour marks the band.

run panel-3model-v122026-09-09cost $0.46mash 0.1.0snapshot ad4b868cde3astatus completed
1 of 10 chapters (Ch 02) were edited after their last judge run. Their scores below are real but outdated, marked stale, and excluded from the Book row average so an old number can’t prop up the headline.
ChapterHumannessVoiceUsefulnessEvidenceDefensibilityNon-redundancy
Book
86
88
61
86
93
87
01 The Shift: From Assistant to Delegate
86
95
47
88
94
100
02 Taste Still Matters When Code Gets Cheapstale
86
85
60
84
94
90
03 Harnesses, Specs, and Codebases Agents Can Actually Use
86
85
66
80
93
85
04 Evals Are the Control System
84
85
61
83
93
85
05 Context Is Infrastructure
86
85
61
90
93
85
06 Runtimes, State, and the Human Control Plane
87
92
65
84
93
85
07 Security, Identity, and High-Stakes Trust
87
95
74
90
92
85
08 Realtime, Voice, and the Cost of Being Interruptible
87
85
65
78
91
85
09 The AI-Native Organization
86
85
65
90
93
85
10 What Endures
88
85
47
90
95
90

Reading the usefulness floor

61
all prose paragraphs
63
substantive core · +1.6

The usefulness judge scores every paragraph for operational density, so it floors the connective tissue every narrative chapter is made of — transitions, scene-setters, recaps, and chapter 10’s reflective register. The substantive core re-averages over the operational paragraphs, setting aside 23 of 532 (4%) that are either markdown headings mis-scored as prose or short bridge sentences the judge’s own rationale names as a transition. The gap is the genre ceiling; the floor that remains is real — prose that could carry a decision, a threshold, or a test and doesn’t yet.

Ship-blockers — unsupported claims

No ship-blockers — every claim defensible against the ledger.

Trend over versions

Humanness-2
Voice-1
Usefulness-15
Evidence+0
Defensibility+4
Non-redundancy+5