LLM Readability Test Protocol

Status: protocol defined, harness written, one subagent pre-test run (tests/readability/2026-10-05-e41258c/notes.md) and two live rounds (tests/readability/2026-10-05-1623155/notes.md, decisions R1 to R8 came from it; tests/readability/2026-10-05-ed37120/notes.md, the four models of R8 on the cheat sheet after R1 to R8). Explain is graded by Claude Code subagents since the second round, and the first round was re-graded by them (decision U4); a hand verdict extends to every sample repeating the sentence it called wrong (U5), the five-point rule names the gating models (U6), and the samples come through the subscription channels unless the owner authorizes the keys (U7). A third round (tests/readability/2026-10-06-3e7c45a/notes.md, Sonnet 5.5 through the Claude Code CLI; gpt-5.5 stopped at 72 samples by the Codex quota) found the cheat sheet's rendering of U2 and U3 wrong and is recorded as a defect round; the library list carries parameter names since (U8). Round 4 (tests/readability/2026-10-06-c696747/notes.md, Sonnet 5.5 on the corrected sheet through the cleaned channel) measures U1 to U3 and U8 on one gating model: Predict 90, Explain 100, Complete 79, Write 40, with no named single argument, rounded() or field-form error left; gpt-5.5 has not run on this sheet (deferred by the owner). The sheet since says when a refined construction takes otherwise (U9), not yet measured. Date: 2026-10-06.

Decision H2 froze the grammar by measurement, not by implementation: before the parser is written, several models must read the cheat sheet and work with the corpus at a correctness rate that justifies the design. Decision V1 (2026-10-06) freezes the grammar by a decision entry instead (V11), after the gap audit's work; the protocol stays as the measurement the owner can ask for. This document fixes the protocol so that the result is comparable across grammar revisions and across models.

What is measured

Four tasks, each scored per program per model. Every prompt contains only the cheat sheet (docs/cheatsheet.md) and the task; no other Renyi material.

Task Prompt Scored by
Predict "Here is a Renyi program and its input. What does it print?" exact match of the predicted output against the reference output
Explain "Explain what this program does in three sentences." a second model grades the explanation against the author's language-neutral description of the program (tests/readability/reference/<program>.explain.txt), blind to which grammar revision produced the program; a second grader grades every explanation too, and a disagreement of more than one point is adjudicated by hand; a hand verdict on a wrong sentence extends to every sample that repeats it (decision U5)
Complete a program with one function body removed, the signature and purpose kept the completed body passes the program's example: and test blocks, judged by hand until M3 and by the VM afterwards
Write a one-paragraph task description the written program, after renyi format, passes the lint, then the same acceptance tests as Complete

The Write task also records the number of lint violations per program, broken down by rule, so that each grammar rule's cost in model errors is visible. renyi format runs before the lint for Complete and Write (decision M5): an agent runs the formatter anyway, so layout problems it fixes are reported but do not fail a sample; the harness keeps a strict tally on request.

Models and settings

At least three models from at least two vendors, including the smallest model the project intends to support (the floor model, decision M7). The gating models are the current large model of each vendor (decision R1: Sonnet 5.5, and gpt-5.5 from the second round, decision R8); the floor model is run and reported as a trend and does not gate the freeze. Temperature 0 where the API allows it, the model's default otherwise. The samples come through the vendor APIs when the owner authorizes the keys for the round, otherwise through the Claude Code CLI (the Claude models) and the Codex CLI (the OpenAI models) on their subscription logins, which the harness drives itself (run --provider claude, run --provider codex; decision U7); the channel, the session and the thinking tokens are recorded with the samples. Five samples per task per program; a task passes for a program when at least four of five samples are correct. Prompts, raw outputs and scores are committed under tests/readability/<date>-<grammar-revision>/.

Acceptance for freezing the grammar

A grammar change that lowers a gating model's passing rate by more than five points is reverted or accompanied by a decision entry explaining why the loss is worth it; the floor model's moves are reported and explained in the round's notes (decision U6).

Regression

The same protocol reruns in CI on every change to docs/cheatsheet.md, to the grammar, or to the corpus, with one sample per task to keep cost low, and with the full five samples before each release.

Known limits

Models trained after the corpus is public may have seen it; the Write task with fresh task descriptions is the control. The Explain grader shares the biases of the grading model; two graders are used when the two disagree by more than one point on a five-point scale.