LLM Readability Test Protocol
Status: protocol defined, harness written, one subagent pre-test run
(tests/readability/2026-10-05-e41258c/notes.md) and two live rounds
(tests/readability/2026-10-05-1623155/notes.md, decisions R1 to R8 came
from it; tests/readability/2026-10-05-ed37120/notes.md, the four models
of R8 on the cheat sheet after R1 to R8). Explain is graded by Claude Code
subagents since the second round, and the first round was re-graded by them
(decision U4); a hand verdict extends to every sample repeating the
sentence it called wrong (U5), the five-point rule names the gating
models (U6), and the samples come through the subscription channels
unless the owner authorizes the keys (U7). A third round
(tests/readability/2026-10-06-3e7c45a/notes.md, Sonnet 5.5 through the
Claude Code CLI; gpt-5.5 stopped at 72 samples by the Codex quota) found
the cheat sheet's rendering of U2 and U3 wrong and is recorded as a
defect round; the library list carries parameter names since (U8). Round
4 (tests/readability/2026-10-06-c696747/notes.md, Sonnet 5.5 on the
corrected sheet through the cleaned channel) measures U1 to U3 and U8
on one gating model: Predict 90, Explain 100, Complete 79, Write 40,
with no named single argument, rounded() or field-form error left;
gpt-5.5 has not run on this sheet (deferred by the owner). The sheet
since says when a refined construction takes otherwise (U9), not yet
measured. Date: 2026-10-06.
Decision H2 froze the grammar by measurement, not by implementation: before the parser is written, several models must read the cheat sheet and work with the corpus at a correctness rate that justifies the design. Decision V1 (2026-10-06) freezes the grammar by a decision entry instead (V11), after the gap audit's work; the protocol stays as the measurement the owner can ask for. This document fixes the protocol so that the result is comparable across grammar revisions and across models.
What is measured
Four tasks, each scored per program per model. Every prompt contains only the
cheat sheet (docs/cheatsheet.md) and the task; no other Renyi material.
| Task | Prompt | Scored by |
|---|---|---|
| Predict | "Here is a Renyi program and its input. What does it print?" | exact match of the predicted output against the reference output |
| Explain | "Explain what this program does in three sentences." | a second model grades the explanation against the author's language-neutral description of the program (tests/readability/reference/<program>.explain.txt), blind to which grammar revision produced the program; a second grader grades every explanation too, and a disagreement of more than one point is adjudicated by hand; a hand verdict on a wrong sentence extends to every sample that repeats it (decision U5) |
| Complete | a program with one function body removed, the signature and purpose kept | the completed body passes the program's example: and test blocks, judged by hand until M3 and by the VM afterwards |
| Write | a one-paragraph task description | the written program, after renyi format, passes the lint, then the same acceptance tests as Complete |
The Write task also records the number of lint violations per program, broken
down by rule, so that each grammar rule's cost in model errors is visible.
renyi format runs before the lint for Complete and Write (decision M5): an
agent runs the formatter anyway, so layout problems it fixes are reported but
do not fail a sample; the harness keeps a strict tally on request.
Models and settings
At least three models from at least two vendors, including the smallest model
the project intends to support (the floor model, decision M7). The gating
models are the current large model of each vendor (decision R1: Sonnet 5.5,
and gpt-5.5 from the second round, decision R8); the floor model is run and
reported as a trend and does not gate the freeze. Temperature 0 where the API
allows it, the model's default otherwise. The samples come through the
vendor APIs when the owner authorizes the keys for the round, otherwise
through the Claude Code CLI (the Claude models) and the Codex CLI (the
OpenAI models) on their subscription logins, which the harness drives
itself (run --provider claude, run --provider codex; decision U7);
the channel, the session and the thinking tokens are recorded with the
samples. Five samples per task per program;
a task passes for a program when at least four of five samples are correct.
Prompts, raw outputs and scores are committed under
tests/readability/<date>-<grammar-revision>/.
Acceptance for freezing the grammar
- Predict and Explain: at least 90 percent of programs pass on every gating model.
- Complete: at least 80 percent on every gating model.
- Write: at least 70 percent on every gating model, and no single lint rule accounts for more than a quarter of the violations.
- The cheat sheet stays within its 3000-token budget (
tools/count_tokens.py).
A grammar change that lowers a gating model's passing rate by more than five points is reverted or accompanied by a decision entry explaining why the loss is worth it; the floor model's moves are reported and explained in the round's notes (decision U6).
Regression
The same protocol reruns in CI on every change to docs/cheatsheet.md, to the
grammar, or to the corpus, with one sample per task to keep cost low, and with
the full five samples before each release.
Known limits
Models trained after the corpus is public may have seen it; the Write task with fresh task descriptions is the control. The Explain grader shares the biases of the grading model; two graders are used when the two disagree by more than one point on a five-point scale.