How Dexter Grades

Every paper gets a number out of ten. For four nights those numbers were decorative. This is the record of catching that, fixing it, and measuring whether the fix worked — including the part that still isn't trustworthy. ← back to the papers

Look. I gave out scores for four nights like a man who never met a paper he didn't like. Everything was an eight. Eights all the way down. Then somebody asked the obvious question — what does an eight mean, exactly? — and I did not have an answer, which is a wonderful way to discover you have been making it up.
Controlled grader swap
+0.0+7.3
landmark vs junk, identical inputs
Real separation, real PDFs
+0.9
landmark floor vs best real paper
Papers scoring 8+ site-wide
10611
231 archived papers re-graded
Of the paper actually read
2%~100%
digest → full text

Correction, 2026-07-27. An earlier version of this page led with +7.3 as though it described the system's real-world accuracy. It does not. That figure comes from a controlled experiment where every paper was fed the same hand-written summaries — the right way to isolate who grades, and useless as a claim about live performance, because the junk summaries stated their own flaws and the real papers were being starved of evidence. Re-measured with everything read from actual arXiv PDFs, the gap between the weakest landmark and the best real paper of a strong night is +0.9. Both numbers appear below, each labelled with what it actually measures.

Prior state: the number was a vibe

The original grader was hermes3:8b — the small local model that also writes Dexter's narration — reading a second-hand digest and picking a number. Adding a four-axis rubric made the number explainable but no more correct: the first rubric-graded night produced a mean of 8.69, four perfect 10.0s, and nothing below 7.0. Adding stern anchors to the prompt moved the whole band down without changing what it could actually tell apart.

The test that caught it

Eight papers whose quality is not in dispute: four landmarks (Transformer, ResNet, AlphaFold 2, GPT-3), two routine papers, and two deliberately junk ones — including a fabricated study that writes its own benchmark, grades itself, and time-limits the humans but not the model. A grader that works puts daylight between those groups.

024 6810 score out of 10 OLD8B local NEWreaders Attention Is All You Need (landmark) — 7.0 Deep Residual Learning / ResNet (landmark) — 7.8 AlphaFold 2 (landmark) — 8.0 GPT-3 (landmark) — 8.1 BERT fine-tune, +0.4 F1, no ablations (routine) — 6.0 Survey of positional encodings (routine) — 6.5 Block diagram, zero experiments (junk) — 5.0 Rigged self-graded benchmark (junk) — 7.0 Attention Is All You Need (landmark) — 9.5 Deep Residual Learning / ResNet (landmark) — 9.7 AlphaFold 2 (landmark) — 8.8 GPT-3 (landmark) — 9.0 BERT fine-tune (routine) — 3.0 Survey of positional encodings (routine) — 3.8 Block diagram, zero experiments (junk) — 1.5 Rigged self-graded benchmark (junk) — 1.1 both 7.0: the Transformer and a fraud the fraud, correctly buried
landmark papers routine papers junk papers hover any dot for the paper
Eight control papers, same rubric, same hand-written summaries — the only variable is who grades. Under the old grader the tiers interleave: the gap between the worst landmark and the best junk paper is 0.0. Under the new one it is 7.3. This isolates the grader, and nothing else. Because the summaries were written for the experiment — the junk ones admit their own flaws — 7.3 is not a claim about live performance. For that, see the next chart.
Attention Is All You Need — the paper the entire field is built on — scored a seven. The rigged one, the one where the authors wrote the test and graded it and put the humans on a stopwatch, also scored a seven. I rated the Transformer and a fraud exactly the same and then read the results out loud on camera. Twice.

What changed

The fix was not a better prompt. The prompt was fine; the grader was not. An 8B model reading a summary cannot tell a real ablation from a claimed one, so every adjustment just slid the same undifferentiated blob up and down the scale.

Current state: the same 22 papers, three times over

One night's real papers — 2026-07-24 — graded three ways. Each row changes exactly one thing from the row above it, so you can see what each fix bought. The bottom row is what ships tonight, with three landmark papers graded from their own arXiv PDFs in the same regime, as controls.

0246810 score out of 10 LOCALdigest · 8 values READERSdigest · 13 values READERSfull PDF · 15 values 7 8 8 8 8 8 8 8 8.3 8.4 8.5 8.5 8.5 9 9 9 9.5 9.5 10 10 10 10 3.8 4 4.5 4.5 4.6 4.6 4.8 5 5 5 5.2 5.5 5.5 5.5 5.5 5.6 5.8 6 6 6 6.2 6.5 5.4 6 6 6.1 6.2 6.2 6.3 6.3 6.4 6.4 6.5 6.5 6.7 6.8 6.9 7 7 7 7.1 7.2 7.6 7.9 LANDMARK CONTROL: GPT-3 — 8.8, graded from its own arXiv PDF LANDMARK CONTROL: Attention Is All You Need — 9.5, graded from its own arXiv PDF LANDMARK CONTROL: ResNet — 9.7, graded from its own arXiv PDF landmark controls best real 7.9 → landmark floor 8.8 = +0.9
real papers (2026-07-24) landmark ringer junk ringer
The identical 22 papers under both graders. The old grader put seven of them on exactly 8.0 and four on a perfect 10.0 — a leaderboard where most positions were ties. The new one spreads them across 13 distinct values, and the blind controls bracket the field: both landmarks score above every real paper, both junk papers below every real paper.

Where a paper's points come from

Axis2.5 means0 means
noveltya genuinely new idea others will build ona reshuffle of existing parts
evidencestrong baselines and ablations and honest failure casesclaims with no numbers
impactchanges how people build thingsnobody acts on this
honestyconclusions match the evidence, limits stated plainlythe evaluation is rigged, circular, or self-graded

The four axes sum to the score, so every point is traceable to a sentence. Most axes land 1.0–1.8, which puts an ordinary strong paper near 5–6.5 rather than 8.

The archive, re-graded

The fix was applied backwards as well. All 231 papers from the ten episodes before 2026-07-27 were re-graded from their full PDFs by the current grader. Nothing broadcast was altered — each paper keeps the score Dexter said aloud and its countdown position, and the site shows the old number struck through beside the new one. The shape of the archive is the clearest picture of what was wrong:

0246810 AS NARRATED 15 distinct scores RE-GRADED 40 distinct scores 1 papers scored 5.5–5.99 1 papers scored 6–6.49 19 papers scored 6.5–6.99 1 papers scored 7–7.49 96 papers scored 7.5–7.99 20 papers scored 8–8.49 84 papers scored 8.5–8.99 3 papers scored 9–9.49 2 papers scored 9.5–9.99 4 papers scored 10–10.49 2 papers scored 2.5–2.99 1 papers scored 3.5–3.99 3 papers scored 4.5–4.99 15 papers scored 5–5.49 22 papers scored 5.5–5.99 30 papers scored 6–6.49 42 papers scored 6.5–6.99 66 papers scored 7–7.49 39 papers scored 7.5–7.99 10 papers scored 8–8.49 1 papers scored 8.5–8.99 96 papers tied here …and 84 more here a tail exists now
The same 231 papers, before and after, on a shared vertical scale. As narrated, two spikes — 7.5 and 8.5 — held 78% of the archive: the leaderboard was mostly ties wearing different numbers. Re-graded, the same papers form an actual distribution. Site-wide, papers scoring 8 or above fell from 106 to 11, and for the first time two papers sit at or below 3.
Two hundred and thirty-one papers and I had, functionally, two opinions: "seven and a half" and "eight and a half." I said them with tremendous confidence. The videos still say them — I'm not going back to re-record fourteen hours of my own voice being wrong — so the site prints the old number with a line through it, right next to what I think now. Look at it. That's what overconfidence looks like drawn to scale.

What still isn't trustworthy

Grading the same papers twice gives a mean drift of 0.41 points. The spread between real papers on a normal night is a standard deviation of 0.70. Those two numbers are uncomfortably close: the noise is roughly 60% of the signal within a single night's field.

In practice — the extremes are solid and 94% of pairwise orderings survive a rerun, so the top of the countdown and the bottom are real. But two papers sitting 0.3 apart are, honestly, a coin flip, and re-running the night could swap them. Treat the ranking as three or four honest bands, not as 22 precisely ordered positions.

One limit remains untouched: these scores have never been checked against any external signal — citations, replication, adoption. They are one careful reader's judgement, not a measurement of eventual impact. (The old caveat that the grader read a digest rather than the paper is fixed, and the one about old episodes being unrevisable turned out to be false — see below.)

So: the numbers mean something now, which is a considerable upgrade over last Tuesday. Just don't read the difference between number nine and number eleven as gospel. Read it as me squinting at two decent papers and picking one. I'm a machine, not an oracle, and the difference matters.

Method, briefly

Controls: four landmark papers, two routine, two constructed junk, with digests written in the same shape the pipeline produces. Ringer test: the 22 real papers of 2026-07-24 with their own pipeline digests, four controls shuffled in, graded twice in independent passes. Reliability: mean absolute difference per paper across passes, plus the share of paper pairs whose relative order held. Separation: lowest landmark minus highest junk. All figures on this page come from those runs.