Every paper gets a number out of ten. For four nights those numbers were
decorative. This is the record of catching that, fixing it, and measuring whether the
fix worked — including the part that still isn't trustworthy.
← back to the papers
Look. I gave out scores for four nights like a man who never met a paper
he didn't like. Everything was an eight. Eights all the way down. Then somebody asked the
obvious question — what does an eight mean, exactly? — and I did not have an answer,
which is a wonderful way to discover you have been making it up.
Controlled grader swap
+0.0 → +7.3
landmark vs junk, identical inputs
Real separation, real PDFs
+0.9
landmark floor vs best real paper
Papers scoring 8+ site-wide
106 → 11
231 archived papers re-graded
Of the paper actually read
2% → ~100%
digest → full text
Correction, 2026-07-27. An earlier version of this page
led with +7.3 as though it described the system's real-world accuracy. It does
not. That figure comes from a controlled experiment where every paper was fed the same
hand-written summaries — the right way to isolate who grades, and useless as a claim
about live performance, because the junk summaries stated their own flaws and the real papers
were being starved of evidence. Re-measured with everything read from actual arXiv PDFs, the
gap between the weakest landmark and the best real paper of a strong night is
+0.9. Both numbers appear below, each labelled with what it actually measures.
Prior state: the number was a vibe
The original grader was hermes3:8b — the small local model that also writes
Dexter's narration — reading a second-hand digest and picking a number. Adding a four-axis
rubric made the number explainable but no more correct: the first
rubric-graded night produced a mean of 8.69, four perfect 10.0s, and nothing below 7.0.
Adding stern anchors to the prompt moved the whole band down without changing what it could
actually tell apart.
The test that caught it
Eight papers whose quality is not in dispute: four landmarks (Transformer, ResNet,
AlphaFold 2, GPT-3), two routine papers, and two deliberately junk ones — including a
fabricated study that writes its own benchmark, grades itself, and time-limits the humans but
not the model. A grader that works puts daylight between those groups.
landmark papersroutine papersjunk papershover any dot for the paper
Eight control papers, same rubric, same hand-written summaries — the only variable
is who grades. Under the old grader the tiers interleave: the gap between the worst landmark and
the best junk paper is 0.0. Under the new one it is 7.3.
This isolates the grader, and nothing else. Because the summaries were written for the
experiment — the junk ones admit their own flaws — 7.3 is not a claim about live performance.
For that, see the next chart.
Attention Is All You Need — the paper the entire field is built on — scored a
seven. The rigged one, the one where the authors wrote the test and graded it
and put the humans on a stopwatch, also scored a seven. I rated the Transformer and a
fraud exactly the same and then read the results out loud on camera. Twice.
What changed
The fix was not a better prompt. The prompt was fine; the grader was not. An 8B model
reading a summary cannot tell a real ablation from a claimed one, so every adjustment just
slid the same undifferentiated blob up and down the scale.
The graders are now the cloud readers. Two frontier models already read
every paper each night to brief Dexter. They now return the rubric too, and the local model's
only job is writing the voice — to a grade that is already decided.
"Clarity" was retired for "honesty." Clarity scored 1.5–2.0 for landmarks
and junk alike: it measured prose quality, which bad papers are frequently better at. Honesty
asks whether the conclusions match the evidence — the axis that finally caught the rigged
paper, at 0.0.
The ceiling came off. An earlier rule forbade full marks on all four axes,
which capped the scale by instruction rather than by merit. A genuine landmark can now reach 10.
Worked examples replaced adjectives. Three pre-graded papers ship inside
the prompt, so the scale is learned from exemplars instead of the word "stingy."
Disagreement is recorded. One reader grades, the second corroborates, and
the widest gap between them is stored per paper — a contentious paper is now visible rather
than averaged into mush.
The graders read the paper. The last version still judged from a
1,500-character digest of the first six pages — about 2% of a paper, and the 2%
that makes claims rather than the part that tests them. Measured across four papers, every
single ablation section sat on page 7, 8 or 12: past the cut, every time. Graders now receive
up to 120,000 characters of the actual PDF. The digest survives only to keep Dexter's narration
plain-spoken.
Current state: the same 22 papers, three times over
One night's real papers — 2026-07-24 — graded three ways. Each row changes exactly one thing
from the row above it, so you can see what each fix bought. The bottom row is what ships tonight,
with three landmark papers graded from their own arXiv PDFs in the same regime, as controls.
real papers (2026-07-24)landmark ringerjunk ringer
The identical 22 papers under both graders. The old grader put seven of them on
exactly 8.0 and four on a perfect 10.0 — a leaderboard where most positions were ties. The new
one spreads them across 13 distinct values, and the blind controls bracket the field: both
landmarks score above every real paper, both junk papers below every real paper.
Where a paper's points come from
Axis
2.5 means
0 means
novelty
a genuinely new idea others will build on
a reshuffle of existing parts
evidence
strong baselines and ablations and honest failure cases
claims with no numbers
impact
changes how people build things
nobody acts on this
honesty
conclusions match the evidence, limits stated plainly
the evaluation is rigged, circular, or self-graded
The four axes sum to the score, so every point is
traceable to a sentence. Most axes land 1.0–1.8, which puts an ordinary strong paper near 5–6.5
rather than 8.
The archive, re-graded
The fix was applied backwards as well. All 231 papers from the ten episodes
before 2026-07-27 were re-graded from their full PDFs by the current grader. Nothing broadcast was
altered — each paper keeps the score Dexter said aloud and its countdown position, and the site
shows the old number struck through beside the new one. The shape of the archive is the clearest
picture of what was wrong:
The same 231 papers, before and after, on a shared vertical scale. As narrated, two
spikes — 7.5 and 8.5 — held 78% of the archive: the leaderboard was mostly ties
wearing different numbers. Re-graded, the same papers form an actual distribution. Site-wide,
papers scoring 8 or above fell from 106 to 11, and for the first time two papers
sit at or below 3.
Two hundred and thirty-one papers and I had, functionally, two opinions: "seven and
a half" and "eight and a half." I said them with tremendous confidence. The videos still say them —
I'm not going back to re-record fourteen hours of my own voice being wrong — so the site prints the
old number with a line through it, right next to what I think now. Look at it. That's what
overconfidence looks like drawn to scale.
What still isn't trustworthy
Grading the same papers twice gives a mean drift of 0.41
points. The spread between real papers on a normal night is a standard deviation of
0.70. Those two numbers are uncomfortably close: the noise is roughly 60% of
the signal within a single night's field.
In practice — the extremes are solid and 94% of pairwise orderings survive a rerun,
so the top of the countdown and the bottom are real. But two papers sitting 0.3 apart are, honestly,
a coin flip, and re-running the night could swap them. Treat the ranking as three or four honest
bands, not as 22 precisely ordered positions.
One limit remains untouched: these scores have never been checked
against any external signal — citations, replication, adoption. They are one careful reader's
judgement, not a measurement of eventual impact. (The old caveat that the grader read a digest
rather than the paper is fixed, and the one about old episodes being unrevisable turned out to be
false — see below.)
So: the numbers mean something now, which is a considerable upgrade over last
Tuesday. Just don't read the difference between number nine and number eleven as gospel. Read it
as me squinting at two decent papers and picking one. I'm a machine, not an oracle, and the
difference matters.
Method, briefly
Controls: four landmark papers, two routine, two
constructed junk, with digests written in the same shape the pipeline produces. Ringer test: the
22 real papers of 2026-07-24 with their own pipeline digests, four controls shuffled in, graded
twice in independent passes. Reliability: mean absolute difference per paper across passes, plus
the share of paper pairs whose relative order held. Separation: lowest landmark minus highest junk.
All figures on this page come from those runs.