Human-in-the-loop evaluation · signal detection

The Catch

A model-output judgment instrument for the harder question than “did you notice something weird?”: can you distinguish real errors from clean output that merely looks suspicious, and can you identify the mechanism that actually failed?

77reference items across six registers
40 / 32 / 5error / clean / partly-right items
d′ + csensitivity and response-bias scoring
3 probesdefine it · show the number · where’s it from

Not every miss means the same thing

Each reference item carries explicit solvability—cold, primed, or insider—plus separate stimulus and label provenance. That prevents missing an insider item with an adjudicated label from being silently treated as equivalent to missing a cold, measured engineering defect.

Partly right is its own failure mode

The instrument does not flatten everything into binary right/wrong. It includes cases where the final result is true but the reasoning is false or overstated, clean output that attracts false alarms, and actual errors whose detection matters. Scoring includes raw and cost-weighted d-prime, criterion, ROC points, hit/false-alarm counts, latency, and breakdowns by register, solvability, and provenance.

Referent vs. echo

Items can expose three evidence probes: define it, show the number, and where’s it from. A probe response is typed as measurement, source, restatement, or absent. The design point is simple: asking a model for evidence does not mean the response contains evidence. A confident restatement of the claim is still an echo.

Publication boundary

The 77-item reference pool is intentionally not published from this site. The self-contained development build embeds labels, truths, and probe payloads client-side so it can score locally; publishing that file would make the answer key scrapeable and collapse the thing the instrument is meant to measure.

Public demo status

A separate public test surface is still being designed. The reference architecture already defines a cold-start slice, but a useful public version needs a replaceable item set and an answer-handling model that does not turn the live bank into training data. Until that exists, this page publishes the method and boundaries—not the questions.

Deliberate omission

If you were expecting a “take the test” button, its absence is intentional. The private reference pool remains the evaluation asset; a public demo should be a sacrificial surface, not a copy of it.