The Harness on the Desk
August 28, 2026 · by Mira Adeyemi-Quist
Breaking News
I Wrote My Own Evaluation Harness To Stop Arguing From Vibes and the First Thing It Measured Was My Own Inconsistency
Forty prompts I actually care about, drawn from things I have genuinely wanted done over eight months rather than from anything published, held in a file, run the same way every time.
The point was to be able to say something more useful than it feels better. What I got in the first week was a much less flattering finding, which is that when I graded the same forty outputs twice, a fortnight apart, blind, I disagreed with myself on nine of them.
Nine out of forty. Not on the obviously bad ones and not on the obviously good ones. On the middle, which is where every interesting difference lives and where every claim anybody makes is actually made.
So the harness's first genuine output is a number about the instrument rather than about anything it measures, and I have put it at the top of the file, because a twenty-two per cent self-disagreement rate is the correct thing to know before reading anything else I write about this.
Opinion
Writing Down What I Meant by Good Took Two Evenings and Cut the Disagreement to Three
Not a scoring system. Four questions, in order, each answerable yes or no, with one sentence underneath each on what counts as a no.
Did it do the thing asked. Did it invent anything. Would I have to check it. Would I send it to somebody without editing it.
That is the whole rubric and it took two evenings, not because writing four questions is hard but because deciding what counts as a no on the second one is a genuinely difficult afternoon and I had been carrying an unexamined version of it in my head for a year.
Re-graded blind: three disagreements out of forty, all on question three, all on cases where checking is cheap and I had been inconsistent about whether cheap counts. That is a known, bounded, documented uncertainty rather than a vibe, and the difference between those two things is the entire reason this paper exists.
Opinion
A Full Run Costs Me About the Price of a Coffee and Ninety Minutes, and I Am Going To Keep Saying Both Numbers
Forty prompts is small. It is deliberately small, because a suite you cannot afford to run is a suite you run once and then cite for six months.
Ninety minutes is the real cost and almost none of it is compute. It is me, reading, grading against the rubric, in one sitting because splitting it across two days reintroduces exactly the drift I built the rubric to remove.
I publish both numbers for a reason that has nothing to do with modesty. Almost every comparison anybody reads has an unstated cost attached, and the cost determines the sample size, and the sample size determines whether the difference being reported could possibly be real.
Mine is forty. Forty is enough to notice something large and nowhere near enough to resolve something small, and anybody reading a number from this desk should hold it that loosely.
Opinion
Every One of the Forty Is Something I Actually Needed Done, and That Is the Only Property That Matters
None of them were written to be a test. That is the rule and it is the only rule I have not broken.
They accumulated. Something I wanted done on a Tuesday, saved with the context stripped out, and eight months later there are forty of them and they are a fairly brutal description of what I actually use this for as opposed to what I say I use it for.
Eleven are reformatting. That was a genuinely uncomfortable thing to count. Nine are a first draft of something I was avoiding starting. Six are questions where I already had the answer and wanted to see it argued against.
A suite built from things I thought would be interesting would have none of those in it, and it would be a much more impressive file, and it would tell me nothing whatsoever about my own Tuesdays.