You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
WS6 — Eval spine: research episodes in Prova, per-role qualification, the measured correction curve
Important
Problem — The above-threshold condition (parent #368) is measured, not assumed — but nothing today measures corrector error rates per role per endpoint. Prova's corpus is pulse-scenario shaped; the qualification battery is endpoint-shaped, not role-shaped; and the correction curve (logical-error vs shots) exists only as hunt fieldnotes.
Approach — Generalize the Prova corpus to research episodes (the recorded QEC hunts become replayable adversarial episodes, including known corrector-discipline traps), extend the qualification battery to score per role, and make the correction curve a first-class reported metric per role per endpoint.
Scope — in: research episode format, per-role battery scoring, correction-curve reporting · out: model fine-tuning (the Specialist consumes this ledger later; it is never a dependency).
Acceptance Criteria
Research episodes (recorded QEC hunts, incl. adversarial corrector-discipline cases) replay through the harness as scored episodes
The qualification battery reports per-role scores; an endpoint's eligibility for a role is measured, dated, and revocable
The correction curve (logical-error vs shots/trials) is reported per role per endpoint and persisted for comparison across runs
The corpus pairing rule extends to research episodes without breaking the existing 29-scenario pulse corpus
Testing Decisions
Extend the corpus reader's pairing-rule tests to research episodes; battery fixtures with mock endpoints of known error rates (the correction curve must recover the planted rates); no parallel eval system — the same corpus serves golden evals and qualification.
Key Decisions
The measured correction curve is the product metric: a flat curve means fix the corrector, not add researchers — the harness reports it rather than burying it.
Adversarial episodes are authored from real failures (surrogate-inflation collapses, refuted claims), never synthetic strawmen.
Constraints & Invariants
Qualification is measured, never assumed (telaio invariant); episode provenance (which hunt, which harness version) is recorded on every score.
Prior Art
The Prova corpus and pairing-rule spec; the telaio qualification battery; the QEC hunt fieldnotes (trial-depth floors, refutation calibration) as episode sources.
Source
Part of #368 · Blocked by #369 · design-of-record: vault spec note spec-20260813-162300-autoresearch-studio
WS6 — Eval spine: research episodes in Prova, per-role qualification, the measured correction curve
Important
Problem — The above-threshold condition (parent #368) is measured, not assumed — but nothing today measures corrector error rates per role per endpoint. Prova's corpus is pulse-scenario shaped; the qualification battery is endpoint-shaped, not role-shaped; and the correction curve (logical-error vs shots) exists only as hunt fieldnotes.
Approach — Generalize the Prova corpus to research episodes (the recorded QEC hunts become replayable adversarial episodes, including known corrector-discipline traps), extend the qualification battery to score per role, and make the correction curve a first-class reported metric per role per endpoint.
Scope — in: research episode format, per-role battery scoring, correction-curve reporting · out: model fine-tuning (the Specialist consumes this ledger later; it is never a dependency).
Acceptance Criteria
Testing Decisions
Extend the corpus reader's pairing-rule tests to research episodes; battery fixtures with mock endpoints of known error rates (the correction curve must recover the planted rates); no parallel eval system — the same corpus serves golden evals and qualification.
Key Decisions
Constraints & Invariants
Prior Art
Source
Part of #368 · Blocked by #369 · design-of-record: vault spec note
spec-20260813-162300-autoresearch-studio