Skip to content

PythiaStage.evaluate scores a trained algorithm against a fabricated all-false truth column (mirrors now-fixed MATLAB #58) #345

Description

@andremun

Description: andremun/InstanceSpace issue #58, filed while reviewing this repo's PR #322, documented that MATLAB's PYTHIAevalMode scores every trained classifier against its test-set truth column unconditionally — including a trained algorithm the test set has no data for, whose truth column is a reconciliation artifact (NaN >= x evaluates to false in MATLAB), not real ground truth. That issue's own text states "the Python side was found to correctly mirror it" at the time — i.e. this port deliberately replicated the bug for parity, per this repo's "MATLAB is the behavioral authority" policy.

MATLAB's issue #58 was closed today (2026-09-26), fixed by andremun/InstanceSpace@c62408f. This port has not been updated to match.

Where the old (now-stale) behavior lives:

  • instancespace/stages/pythia.py::PythiaStage.evaluate loops over every non-None fitted classifier slot and scores it against y_true, with no way to know whether that column reflects real test data or _build_test_algo_matrix's reconciled all-NaN/all-False placeholder.
  • instancespace/instance_space.py::InstanceSpace._build_test_algo_matrix already computes a has_ground_truth mask (used for compute_binary_performance's widened comparison) but _explore_evaluate (instance_space.py:1215) discards it: y_raw_test, _ = self._build_test_algo_matrix(...).
  • The docstrings at instance_space.py:1183-1187 and pythia.py:472-479 both currently document the old MATLAB behavior as intentional ("exactly as MATLAB does").

Evidence this is real, not theoretical: a report-only CI job (matlab-oracle-verify.yml, added on PR #343) ran the real MATLAB exporter against andremun/InstanceSpace's current master and, after separating platform/run noise from real code-driven drift (control run pinned at the old gold commit vs. current master), found exactly 8 files whose content is attributable to actual MATLAB code changes since the pinned commit — all 8 are PYTHIA's own hyperparameter/selection/eval outputs for the legacy_svm and trace3_pythia_skip variants (build_data/pythia/legacy_svm/outputs/{hyperparameters,raw_metrics,selection0,selection1,summary,ysub}.csv, explore_data/pythia/legacy_svm/outputs/eval_summary.csv, explore_data/pythia/trace3_pythia_skip/outputs/eval_summary.csv). Full breakdown on #344.

Suggested fix, mirroring MATLAB's own resolution:

  • Thread has_ground_truth through to PythiaEvaluateInput (new field).
  • In PythiaStage.evaluate, skip the confusion-matrix computation for a trained-algorithm index where has_ground_truth[index] is False, leaving its accuracy/precision/recall as NaN and its cvcmat row as zero — the same treatment already given to an empty (None) classifier slot.
  • Update instance_space.py::_explore_evaluate to pass the real mask instead of discarding it.

Safety check already done: no current tests/fixtures/matlab/current-based parity test reads the specific 8 affected files for a PythiaStage.evaluate comparison — test_current_matlab_trace_parity.py's similarly-named raw_metrics.csv/eval_summary.csv reads are under build_data/trace//explore_data/trace/, a different stage directory from the affected pythia/ files. test_current_matlab_stages.py's one PythiaEvaluateInput call site is the trace3_pythia_skip variant, whose classifier slots are already all None (training was skipped), so the fix doesn't change that test's outcome. Fixing this code should not require touching tests/fixtures/matlab/current or its pinned manifest.

Compatibility tag: [Behavior-changing] — InstanceSpace.explore()'s reported accuracy/precision/recall/cvcmat for a trained algorithm absent from a test set's metadata changes from a fabricated score to NaN/zero-row, matching corrected MATLAB behavior.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions