Skip to content

fix: refresh MATLAB fixture bundle (98a01ac → fdad7a43) and PYTHIA eval scoring (#345) - #346

Open
andremun wants to merge 6 commits into
mainfrom
matlab-fixture-refresh-2026-09-26
Open

andremun wants to merge 6 commits into
mainfrom
matlab-fixture-refresh-2026-09-26

Conversation

@andremun

@andremun andremun commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Compatibility tag: [Behavior-changing]. Closes #345. Addresses the fixture-staleness half of #344 (the CI gate check itself — validation-tests.yml's oracle-SHA check and manifest.json's pin — isn't touched by this PR and needs its own follow-up once this merges).

What this does

Replaces the "verified v2" MATLAB oracle at tests/fixtures/matlab/current/ — generated under the pinned commit 98a01ac0513c0dd0f8a9bd91ed2926c871334d7b (v0.9.1, on macOS) — with a fresh export from andremun/InstanceSpace's current master, fdad7a43a12a403c4b20e4e98bbf8a6e56538e2b, generated on Linux. This is a real MATLAB run, not a synthetic or reasoned-about bundle: matlab-fixture-publish.yml (new, one-off workflow_dispatch tool, not for routine use, included in this PR) ran the actual tests/matlab_export/pyis_export_reference_data.m exporter under real MATLAB R2026a on a GitHub-hosted runner and pushed the raw 423-file result directly to this branch by git (the session's sandbox can reach github.com but not the Actions artifact blob-storage host, so a direct push was the only way to get the export out of CI and into something reviewable).

Why this needed care, not just a data swap

The drift wasn't what it first looked like. Naively diffing the old oracle against a fresh master export showed 190 of 423 files changed. A control run — re-exporting at the exact old gold commit, on this same Linux CI runner — still showed 182 of those 190 files different, with zero MATLAB code difference. That's platform/run noise: the old oracle was generated on macOS; SIFTED's genetic algorithm and PILOT's iterative solver are both sensitive to BLAS/LAPACK-level floating-point differences across platforms.

Correction, posted 2026-09-26: this PR originally claimed "only 8 files are real, code-attributable drift." Two further independent re-exports at this PR's own new gold commit reproduced the same 182-file diff against the just-committed oracle, which exposed a flaw in how that 8-file figure was isolated (see the full correction comment on this PR and on #344). Corrected finding: ~175 of those 182 files are unstable at both the old and new gold commit — inherent run-to-run non-determinism, not narrowed by generating on Linux. Only 7 files show genuinely new instability introduced between the two commits, all under pythia/legacy_svm outputs (hyperparameters, raw_metrics, selection0, selection1, summary, ysub, plus explore's eval_summary), plausibly tied to #58's changed scoring logic but confirmed only at file-identity level, not value level. This does not change the fixture content itself, only the drift characterization — see the correction comment for the full breakdown.

The #58 fix turned out broader than first modeled. #345's original port only changed "trained algorithm absent from test-set ground truth" to NaN. Reading andremun/InstanceSpace@master's actual current core/PYTHIA.m (not just the issue text) while debugging a test failure against the new fixtures showed the real fix also changed an empty (untrained) classifier slot from "zero accuracy, undefined precision/recall" to fully NaN — MATLAB now unifies "no classifier" and "no ground truth" into one "not scored" case. Corrected PythiaStage.evaluate and its test to match. A subsequent Copilot review round found this still wasn't complete — MATLAB masks per-instance (observed = ~isnan(Y(:,ii))), not per whole algorithm column — so PythiaEvaluateInput.observed was reworked to a full (n_instances, n_trained) mask, with a new end-to-end test through _explore_evaluate proving per-instance masking (not just the denominator) actually changes scoring.

Two test files needed non-mechanical updates, not just fixture-path swaps: a literal old-commit/platform assertion (test_current_bundle_is_verified_r2026a_source), and two pinned dicts in test_current_matlab_trace_parity.py (_BOUNDARY_AMBIGUITIES, _BOUNDARY_SUMMARY_VALUES) that recorded exactly one floating-point boundary-ambiguous point per TRACE3 variant under the old oracle — the refreshed oracle's real geometry has none for either variant (confirmed by the membership test's own actual-vs-expected diff going from 1 point to 0, not assumed or guessed).

What did NOT need to change

  • tools/fixture_provenance.py's _REFERENCE_V2_EXPORTER_SHA256 and _CANONICAL_DATASET_SHA256 — same exporter script, same input dataset.
  • tests/fixture_inventory.json — the file path set is unchanged (still 423 files at the same paths), and the inventory classifies trust by path, not content hash.

Verification

  • tools/fixture_provenance.py verify tests/fixtures/matlab/current passes (matlab-verified, 423 files).
  • poe check (black, ruff, mypy --strict, docs) passes.
  • Full pytest suite: 1049 passed (pytest --collect-only -q confirms 1049 collected), 92.05% branch coverage (above the 85% gate).
  • Every change to a pinned/expected value in the test suite is backed by re-reading either MATLAB's actual current source or the fresh fixture's own data — none guessed from the issue text or the old value's shape.

Dependency note

This branch cherry-picks #345's commit (f9def48 on #343) because the refreshed fixture data requires it — without it, test_current_matlab_pythia_skip_oracle fails. If #343 merges first, this PR's diff for those files will simplify to nothing extra. If this PR merges first, #343's own commit will no-op cleanly on rebase (identical patch).

Still open after this merges

  • validation-tests.yml's "Check MATLAB oracle is current" step and manifest.json's pin still need bumping to fdad7a43a12a403c4b20e4e98bbf8a6e56538e2b — deliberately left to a separate follow-up PR, so the fixture-content change (this PR, reviewable on its numeric merits) and the CI-gate change (mechanical, once this is approved) aren't bundled into one review.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25


Generated by Claude Code

actions-user and others added 3 commits September 26, 2026 08:18
Raw output of pyis_export_reference_data.m against InstanceSpace@master,
pushed as-is by matlab-fixture-publish.yml. tools/fixture_provenance.py's
pinned identity constants, tests/fixture_inventory.json, and the test
suite have NOT been updated yet -- this commit is raw material for
that follow-up work, not a reviewed fixture bundle on its own.
…nt from test data (#345)

[Behavior-changing] InstanceSpace.explore()'s PYTHIA evaluation used to
score every trained classifier against its reconciled test-set column
unconditionally, including a trained algorithm the test set simply has
no data for -- whose "truth" column is an all-false reconciliation
artifact, not a real measurement. This was a deliberate bit-for-bit
mirror of a real MATLAB defect (andremun/InstanceSpace#58), recorded as
intentional at the time. MATLAB fixed #58 today (c62408f); this ports
the same fix.

PythiaEvaluateInput gains a required has_ground_truth field (one entry
per trained classifier). PythiaStage.evaluate now skips the confusion-
matrix computation for a trained index with no ground truth, leaving its
accuracy/precision/recall as NaN and its confusion row zero -- the same
treatment already given to a test-only algorithm with no trained-model
slot. InstanceSpace._explore_evaluate now threads the mask
_build_test_algo_matrix already computed instead of discarding it.

No fixture or manifest changes needed: no current tests/fixtures/matlab/
current parity test reads the 8 files this fix's real-world effect maps
to (isolated via matlab-oracle-verify.yml's gold-commit control run) for
a PythiaStage.evaluate comparison. Full pytest suite: 1047 passed,
unchanged. Rewrote the one test asserting the old fabricated-score
behavior into the new skip behavior, and added a companion test proving
the fix distinguishes "no test data" from "observed and genuinely bad
everywhere" -- both produce an all-false column, but only the former is
missing data.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25
… scoring

[Behavior-changing] Pins tools/fixture_provenance.py's gold identity
(_GOLD_MATLAB_COMMIT, _VERIFIED_V2_CONTENT_ROOT_SHA256) to the raw bundle
matlab-fixture-publish.yml pushed in the previous commit, replacing the
stale 98a01ac.../macOS oracle with fdad7a43.../Linux.

Reading andremun/InstanceSpace@master's actual current core/PYTHIA.m
(not just the #58 issue text) showed the real fix is broader than
f9def48 modeled: an untrained (empty) classifier slot now also gets
fully NaN accuracy/precision/recall, not the old "zero accuracy,
undefined rates" treatment -- MATLAB unifies "no classifier" and "no
ground truth" into the same "not scored" case. Corrected
PythiaStage.evaluate's two internal loops and the one test asserting
the old empty-slot behavior to match.

Also updates: the literal old-commit assertion in
test_current_bundle_is_verified_r2026a_source; and
test_current_matlab_trace_parity.py's _BOUNDARY_AMBIGUITIES/
_BOUNDARY_SUMMARY_VALUES, which pinned the one boundary-ambiguous point
each trace3 variant had under the old oracle -- the refreshed oracle's
real geometry has none for either variant, confirmed by the membership
test's own actual-vs-expected diff, not assumed.

tests/fixture_inventory.json needed no change (same file paths, content
classified by path not hash). Full verification: poe check and the
complete pytest suite (1047 passed, 92.05% coverage) against the
refreshed bundle.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25
Copilot AI lite review requested due to automatic review settings September 26, 2026 08:53

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

PYTHIA still needs correct confusion-row and per-instance masking behavior, plus focused end-to-end coverage.

Review effort: Lite
Findings: 2 Medium severity · 1 Low severity

Open (3)
What changed in this PR

Refreshes the MATLAB oracle fixtures to commit fdad7a43 and aligns PYTHIA scoring with MATLAB’s missing-ground-truth behavior.

Changes:

  • Refreshes provenance metadata and generated MATLAB fixtures.
  • Updates PYTHIA evaluation and related tests.
  • Refreshes TRACE parity expectations.
File Reviewed change
tools/​fixture_provenance.py Updates oracle provenance.
tests/​test_explore_evaluate.py Updates PYTHIA evaluation tests.
tests/​test_current_matlab_trace_parity.py Refreshes TRACE parity expectations.
tests/​test_current_matlab_stages.py Updates oracle metadata expectations.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​trace3_pythia_skip/​outputs/​membership.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​trace3_pythia_skip/​outputs/​eval_summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​trace3_default/​outputs/​membership.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​trace3_default/​outputs/​eval_summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​pilot_standard_analytic_3d/​outputs/​eval_summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​trace/​legacy_svm/​outputs/​eval_summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​pythia/​trace3_pythia_skip/​outputs/​eval_summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​explore_data/​pilot/​pilot_standard_numerical_3d_x0/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
tests/​fixtures/​matlab/​current/​explore_data/​pilot/​pilot_standard_numerical_3d_precalc/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
tests/​fixtures/​matlab/​current/​explore_data/​pilot/​pilot_standard_analytic_3d/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
tests/​fixtures/​matlab/​current/​explore_data/​pilot/​pilot_pls_3d_grouped/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
tests/​fixtures/​matlab/​current/​explore_data/​pilot/​pilot_pls_2d/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_RBF_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_poly_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_NB.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_LDA.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_KNN.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​good_CART.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_poly_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_NB.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_LDA.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_L_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_KNN.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_J48.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_pythia_skip/​outputs/​best_CART.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_default/​outputs/​raw_metrics.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_default/​outputs/​good_J48.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_default/​outputs/​best_RandF.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​trace3_default/​outputs/​best_J48.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​good_poly_SVM_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​good_poly_SVM_alpha_spectrum.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​good_NB_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​good_KNN_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_RBF_SVM_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_RBF_SVM_alpha_spectrum.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_NB_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_NB_alpha_spectrum.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_LDA_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_LDA_alpha_spectrum.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_KNN_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_KNN_alpha_spectrum.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​pilot_standard_analytic_3d/​outputs/​best_CART_vertices.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​summary.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_RBF_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_poly_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_NB.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_LDA.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_KNN.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_J48.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​good_CART.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_RBF_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_NB.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_LDA.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_L_SVM.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_KNN.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​trace/​legacy_svm/​outputs/​best_CART.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​sifted/​default/​outputs/​correlation_rho.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​trace3_pythia_skip/​outputs/​normalization_sigma.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​trace3_pythia_skip/​outputs/​normalization_mu.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​trace3_default/​outputs/​normalization_sigma.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​trace3_default/​outputs/​normalization_mu.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​legacy_svm/​outputs/​normalization_sigma.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pythia/​legacy_svm/​outputs/​normalization_mu.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​viewpoint_angles.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​viewpoint_a.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_perf.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_error.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_eoptim.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_x0/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​viewpoint_angles.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​viewpoint_a.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​pilot_error.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_numerical_3d_precalc/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​viewpoint_angles.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​viewpoint_a.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_standard_analytic_3d/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​viewpoint_angles.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​viewpoint_a.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_3d_grouped/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_2d/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_2d/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_2d/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​pilot_pls_2d/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_r2.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_perf.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_matrix.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_error.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_eoptim.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_c.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_b.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​pilot/​default/​outputs/​pilot_a_raw.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​cloister/​default/​outputs/​z_edge.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​cloister/​default/​outputs/​z_ecorr.csv Refreshes MATLAB fixture output.
tests/​fixtures/​matlab/​current/​build_data/​cloister/​default/​inputs/​projection_a.csv Refreshes MATLAB fixture input.
instancespace/​stages/​pythia.py Updates missing-ground-truth scoring.
instancespace/​instance_space.py Threads ground-truth metadata into evaluation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread instancespace/instance_space.py Outdated
Comment thread instancespace/stages/pythia.py Outdated
Comment thread tools/fixture_provenance.py
[Additive] Same file as on PR #343. Added here specifically to answer a
direct question: does the refreshed bundle actually reproduce on a fresh
run now that generation is Linux-to-Linux (no more macOS-vs-Linux noise)?
Running this on the current head diffs a fresh master export against the
new tests/fixtures/matlab/current (fdad7a43...) committed in this PR,
which InstanceSpace's master still matches exactly at push time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25
@codecov-commenter

codecov-commenter commented Sep 26, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 90.55%. Comparing base (67753e2) to head (5582b50).
⚠️ Report is 7 commits behind head on main.
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #346      +/-   ##
==========================================
- Coverage   90.60%   90.55%   -0.05%     
==========================================
  Files          26       26              
  Lines        5724     5730       +6     
  Branches      700      701       +1     
==========================================
+ Hits         5186     5189       +3     
- Misses        349      351       +2     
- Partials      189      190       +1     

see 12 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

[Behavior-changing] Addresses Copilot review findings on PR #346.

PythiaStage.evaluate's ground-truth mask was per-algorithm-column
(has_ground_truth: shape (n_trained,)), collapsing "some test instances
observed, others not" into an all-or-nothing decision. MATLAB's actual
#58 fix masks per instance: core/PYTHIA.m computes
observed = ~isnan(Y(:,ii)) and scores only those rows, with accuracy's
denominator being the observed count (tp+tn+fp+fn), not the total
instance count. A trained algorithm with partial test-set coverage was
still having its unobserved rows counted as fabricated true/false
negatives, and its accuracy denominator was always n_instances even when
only some instances were observed.

PythiaEvaluateInput.has_ground_truth (shape (n_trained,)) is now
.observed (shape (n_instances, n_trained)); _build_test_algo_matrix
computes it as ~isnan(y_raw_test[:, :n_trained]) instead of a whole-column
boolean. PythiaStage.evaluate masks each algorithm's confusion count to
its observed rows and uses the observed count as accuracy's denominator.

New tests: test_pythia_evaluate_masks_per_instance_missing_observations
(a trained algorithm observed on 2 of 3 test instances -- confirms the
unobserved instance isn't scored and the denominator is 2, not 3, not
just that the result happens to still be 0) and
test_explore_evaluate_trained_algorithm_missing_from_test_set (exercises
_explore_evaluate/_build_test_algo_matrix end-to-end, not
PythiaStage.evaluate directly, so a regression that discards or replaces
the mask with all-True would be caught here even though every
evaluate()-level test would still pass).

Also updates stale documentation Copilot flagged as still describing the
old fabricated-score behavior or the pre-refresh oracle identity:
tests/README.md, tests/matlab_export/README.md, docs/test_data_audit.md,
docs/stage_inference_architecture.md, docs/python_implementation_pathways.md.

Verification: poe check and the full pytest suite (1049 passed, 92.05%
coverage) pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Reconcile the documented test counts, export workflow provenance, and current oracle references before approval.

Review effort: Lite
Findings: 3 Low severity

Open (3)
Resolved since last review (3)

Comment thread tests/README.md
Comment thread tests/matlab_export/README.md
Comment thread tests/matlab_export/README.md

Copy link
Copy Markdown
Owner Author

Correction to this PR's description — the "only 8 files are real, code-attributable drift" claim in the "Why this needed care" section is wrong.

After this fixture refresh was committed, I re-ran matlab-oracle-verify.yml twice more, independently, at this exact same gold commit (fdad7a4). Both reproduced an identical 182-file diff against the committed oracle — not random noise, but a stable discrepancy I hadn't accounted for. Comparing that file list against the old control run's own file list (rather than just against the raw 190/182 counts) shows:

  • ~175 of the 182 files are unstable at both the old and new gold commit — inherent run-to-run non-determinism (SIFTED's GA, PILOT's iterative solver, and everything downstream: TRACE geometry, PYTHIA normalization/probabilities), not narrowed by generating on Linux as the PR description currently claims.
  • Only 7 files are genuinely new instability introduced between the two commits: build_data/pythia/legacy_svm/outputs/{hyperparameters,raw_metrics,selection0,selection1,summary,ysub}.csv and explore_data/pythia/legacy_svm/outputs/eval_summary.csv. These plausibly trace to test automerge #58's changed selection/scoring logic (legacy_svm exercises missing-test-data scoring), but that's confirmed only at file-identity level — this sandbox's proxy blocks the Actions blob-storage host, so I can't pull the raw CSVs for an actual value-level diff.
  • 6 files that were unstable under the old-commit control are now exact matches — consistent with the already-landed pointwise_covers() fix (Codex/open issue big rocks #319), not test automerge #58.

None of this changes the fixture content itself (tests/fixtures/matlab/current/ is still the best available real-MATLAB snapshot at fdad7a4, and the PYTHIA eval-scoring fix in this PR is still correct and verified against MATLAB's actual source). It changes the characterization of the drift and, more importantly, rules out the "automate the gate with a raw file-diff" idea discussed on #344 — the noise floor is too large relative to the 423-file total for file-identity diffing to be a safe pass/fail signal.

Filed the same correction on #344. Leaving the PR description's original paragraph as-is rather than silently editing it, per this repo's own conventions on visible corrections.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

verify_against_real_matlab check (failure, job-level continue-on-error: true): this is the same 182-file drift already covered in the correction comment above — this workflow is explicitly report-only by design (see its own header comment) precisely because it's the exporter's first run outside a human's local MATLAB, and it isn't wired to block anything yet. No fix needed in this PR: the failure is the expected report from comparing this PR's own head commit's fresh export against the just-committed oracle, which is the same stable (not flaky) 182-file discrepancy already analyzed. Re-running it would reproduce the same result, not resolve anything, so I'm not spending a re-run on it.

Also pushing three fixes for Copilot's Low-severity findings on this review round:

  1. matlab-fixture-publish.yml was referenced by tests/matlab_export/README.md but missing from this branch (it existed on a sibling branch/PR) — added it here.
  2. PR description's "1047 passed" reconciled to the actual pytest --collect-only count, 1049 (two tests were added by the earlier per-instance-masking fix on this same branch).
  3. The roadmap's "exactly 8 files are code-attributable" figure (not itself a Copilot finding, but the same area) is corrected by the reproducibility comment above; leaving the historical roadmap entries as-is (append-only convention) and adding a new dated entry instead.

Generated by Claude Code

…ion, missing workflow)

[Additive]. Adds matlab-fixture-publish.yml to this branch (referenced by
tests/matlab_export/README.md's provenance section but missing here; the
self-guarded push trigger is dropped since GitHub already registered the
workflow from its first use on PR #343). Reconciles the PR description's
stale "1047 passed" against the actual 1049 pytest-collected count.
Appends a roadmap v1.78 entry correcting v1.76/v1.77's "8 files are
code-attributable" claim, per two further independent same-commit
re-exports that reproduced an identical 182-file diff and exposed the
comparison error - full detail posted as PR/issue comments rather than
edited into the append-only historical rows.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cr45PPFZJorxbj4dcgVM25

Copy link
Copy Markdown
Owner Author

Posted a value-level correction to my earlier hash-based drift analysis on #344 — same finding applies here: most of the refreshed oracle's changed files (160/183) are within precision, but 18 show real magnitude changes (SVM probability outputs differing up to 40%, a TRACE membership flag flipping, PILOT's alpha differing 32%), not floating-point noise. This doesn't change what's committed in this PR — the new oracle is still the best available real-MATLAB snapshot — but it means the actual behavioral delta between the old and new oracle is larger and more real than "mostly noise" suggested. Full breakdown on #344.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Posted a further finding on #344 that overturns my previous comment here: the platform-isolation experiment shows the "large magnitude" files aren't platform-driven or code-driven — they're inherently non-reproducible run-to-run (likely unseeded randomness in legacy_svm's LIBSVM training and PILOT's multi-restart solver), independent of both variables. Doesn't change anything committed in this PR; the new oracle is still the right thing to have merged. Full detail on #344.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Correction to my last comment's guessed mechanism, posted on #344: it's not LIBSVM (confirmed deprecated and unused for training in current MATLAB — fitcsvm is used instead), it's most likely MATLAB's fitSVMPosterior posterior-probability calibration step, which runs its own internal cross-validation after the seeded fitcsvm call. Full detail on #344.


Generated by Claude Code

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PythiaStage.evaluate scores a trained algorithm against a fabricated all-false truth column (mirrors now-fixed MATLAB #58)

5 participants