Skip to content

WS6 — Eval spine: research episodes in Prova, per-role qualification, the measured correction curve #374

Description

@aarontrowbridge

WS6 — Eval spine: research episodes in Prova, per-role qualification, the measured correction curve

Important

Problem — The above-threshold condition (parent #368) is measured, not assumed — but nothing today measures corrector error rates per role per endpoint. Prova's corpus is pulse-scenario shaped; the qualification battery is endpoint-shaped, not role-shaped; and the correction curve (logical-error vs shots) exists only as hunt fieldnotes.

Approach — Generalize the Prova corpus to research episodes (the recorded QEC hunts become replayable adversarial episodes, including known corrector-discipline traps), extend the qualification battery to score per role, and make the correction curve a first-class reported metric per role per endpoint.

Scope — in: research episode format, per-role battery scoring, correction-curve reporting · out: model fine-tuning (the Specialist consumes this ledger later; it is never a dependency).

Acceptance Criteria

  • Research episodes (recorded QEC hunts, incl. adversarial corrector-discipline cases) replay through the harness as scored episodes
  • The qualification battery reports per-role scores; an endpoint's eligibility for a role is measured, dated, and revocable
  • The correction curve (logical-error vs shots/trials) is reported per role per endpoint and persisted for comparison across runs
  • The corpus pairing rule extends to research episodes without breaking the existing 29-scenario pulse corpus

Testing Decisions

Extend the corpus reader's pairing-rule tests to research episodes; battery fixtures with mock endpoints of known error rates (the correction curve must recover the planted rates); no parallel eval system — the same corpus serves golden evals and qualification.

Key Decisions

  • The measured correction curve is the product metric: a flat curve means fix the corrector, not add researchers — the harness reports it rather than burying it.
  • Adversarial episodes are authored from real failures (surrogate-inflation collapses, refuted claims), never synthetic strawmen.

Constraints & Invariants

  • Qualification is measured, never assumed (telaio invariant); episode provenance (which hunt, which harness version) is recorded on every score.

Prior Art

  • The Prova corpus and pairing-rule spec; the telaio qualification battery; the QEC hunt fieldnotes (trial-depth floors, refutation calibration) as episode sources.

Source

Part of #368 · Blocked by #369 · design-of-record: vault spec note spec-20260813-162300-autoresearch-studio

Metadata

Metadata

Assignees

No one assigned

    Labels

    afkImplement + merge unattended — tests decide green

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions