diff --git a/docs/classifiers.md b/docs/classifiers.md index 72e461c..24dfecd 100644 --- a/docs/classifiers.md +++ b/docs/classifiers.md @@ -114,13 +114,11 @@ Mathematics, 9(23), 2021. This is the only Keras port here, following the authors, so it needs `tensorflow` rather than `torch`. Both are in the `deep-learning` extra. -### Tuning, and what we report - -**The XCM results in this repository follow the authors' protocol.** Section 4.3 sets +The XCM results in this repository follow the authors' tuning protocol. Section 4.3 sets `window_size` and `batch_size` per dataset "by grid search based on the best average accuracy following a stratified 5-fold cross-validation on the training set", over windows {0.2, 0.4, 0.6, 0.8, 1.0} and batches {1, 8, 32}. Selection never touches the -test data, so the published figures are tuned but not leaked, and neither are ours. +test data. The reported run searches the window on that grid and holds batch size at 32. That is the one departure, and it is a cost decision rather than a modelling one: batch 1 takes @@ -161,12 +159,29 @@ reproduce them, and expect the memory cost to follow. ## Notes on the ports -All three wrappers take aeon's ``numpy3D`` collections, shape +All classifiers in this package take aeon's ``numpy3D`` collections, shape ``(n_cases, n_channels, n_timepoints)``, and transpose internally where the original -network expects a different layout. Each holds out ``validation_size`` of the training -data inside ``fit`` and restores the best epoch, so no external validation split is -required. The networks require ``torch``, which is not a hard dependency of this -package: install it with ``pip install aeon-multiverse[deep-learning]``. +network expects a different layout. Validation and epoch selection differ by port: + +* ConvTran and PatchMTSC split off ``validation_size`` inside ``fit`` and restore the + epoch with the lowest validation loss. With ``validation_size=0`` they select on + training loss instead. +* TimesNet splits internally and restores the epoch with the best validation accuracy + when a nonzero split is feasible. With ``validation_size=0`` (or fewer than two + cases) it selects on training loss instead. +* DisjointCNN passes a sampled ``validation_size`` set to Keras, but samples with + replacement from the training collection and leaves those cases in training. It + restores the best monitored weights, but this is not a held-out validation split. +* XCM has no validation split or best-epoch restoration: it trains for a fixed number + of epochs. Its optional cross-validation selects hyperparameters, not an epoch. +* TimesURL and TS2Vec pretrain on the full training collection and fit their probe on + the resulting training representations; neither has an internal validation split or + best-epoch restoration. + +Thus no external validation split is required by these wrappers, but the same +validation procedure does not apply to every classifier. The networks require +``torch``, which is not a hard dependency of this package: install it with +``pip install aeon-multiverse[deep-learning]``. ### What was ported, and from where @@ -216,7 +231,7 @@ faithful transcription. | ``probability=True`` on the SVM probe | TS2Vec | The authors' grid sets it False, which leaves an ``SVC`` unable to produce probability estimates. aeon classifiers must implement ``predict_proba``, so it is enabled, adding Platt scaling fitted by internal cross-validation on the training data | | Layer imports taken from ``tensorflow.keras.layers`` | XCM | The original imports ``Conv1D`` and ``Conv2D`` from ``keras.layers.convolutional``, a path removed in Keras 3. The layers and their arguments are unchanged | | Kernel length floored at one point | XCM | The original computes ``int(window_size * n)``, which is zero for series shorter than five points and builds an invalid layer | -| Validation split moved inside ``fit`` | all three | The originals split train/validation outside the model, which risks leakage between train and test. TSLib is explicit about it: ``exp_classification.py`` sets ``vali_data = self._get_data(flag='TEST')``, so it selects the retained epoch on the test set. See the note at the top of this page | +| Validation split moved inside ``fit`` | ConvTran, PatchMTSC, TimesNet | The originals split train/validation outside the model, which risks leakage between train and test. TSLib is explicit about it: ``exp_classification.py`` sets ``vali_data = self._get_data(flag='TEST')``, so it selects the retained epoch on the test set. See the note at the top of this page | | Test data scaled with training statistics | TimesNet | TSLib fits its normaliser separately per split, so its test set is scaled by its own statistics. The port fits on train and applies to test | Note that PatchMTSC's two graph blocks are, in the original, distinguished only by a diff --git a/docs/datasets.md b/docs/datasets.md index 46b2877..3fe6424 100644 --- a/docs/datasets.md +++ b/docs/datasets.md @@ -57,7 +57,7 @@ Unequal length datasets are stored in a list of 2D numpy arrays. You can control whether to load the equal length version with the parameter ``load_equal_length``. ```python -X,y = load_classification("JapaneseVowels", load_equal_length = False) # Unequal length example +X,y = load_classification("JapaneseVowels", load_equal_length = True) # Unequal loaded by default from 1.5 ``` Imputed missing value versions can be loaded with the argument ``load_no_missing``. You can download whole archives from zenodo or in code @@ -69,9 +69,34 @@ download_archive(archive="UEA", extract_path="C:\\Temp\\") ``` Currently should be one of "EEG","UCR","UEA","Imbalanced","TSR", "Unequal". See -``aeon`` documentation for more details. There are lists of datasets in aeon and a +``aeon`` documentation for more details. There are lists of datasets in aeon and a dictionary of all zenodo keys. ```python from aeon.datasets.tsc_datasets import multiverse_core, multiverse2026, eeg2026, tsc_zenodo ``` + +## Dataset collections used here + +The project uses the archive collections exposed by aeon: + +```python +from aeon.datasets.tsc_datasets import multiverse_core, multiverse2026, eeg2026 + +print(len(multiverse_core)) # 66 +print(len(multiverse2026)) # 133 +print(len(eeg2026)) # 28 +``` + +`multiverse2026` is the full Multiverse collection. `multiverse_core` is the smaller +benchmark subset: it is more balanced across applications, removes overly similar, +very simple and zero-information datasets, and has a useful spread of dataset sizes and +series lengths. The current benchmark results use this core list unless stated +otherwise. + +The EEG collection is a separate classification archive used for EEG-specific +experiments. It is based on [aeon-neuro](https://github.com/aeon-toolkit/aeon-neuro). + +These lists are Python collections of dataset names, so they can be passed directly to +experiment or result-loading code. The underlying dataset files are still downloaded +and cached using `load_classification`, as described above. diff --git a/docs/evaluation.md b/docs/evaluation.md index 364699b..258bde2 100644 --- a/docs/evaluation.md +++ b/docs/evaluation.md @@ -1,11 +1,206 @@ -# Experimental Protocols +# Experimental protocols -There are many variations on how people structure experiments and a range of metrics -used in comparison. +This repository uses [`tsml-eval`](https://github.com/time-series-machine-learning/tsml-eval) +to run classification experiments and store predictions. `tsml-eval` is the evaluation +toolkit used by the time-series machine-learning projects around aeon. It loads the +train and test files for a dataset, fits an aeon-compatible classifier, measures the +run, writes predictions and probabilities, and provides utilities for comparing +classifiers over datasets and resamples. -The results we present are for the moment the simplest: we do all -training/validation on the default train split, and evaluate once on the test set. +The basic protocol here is deliberately simple: fit on the supplied training split, +then evaluate once on the supplied test split. Resample `0` means the archive's default +train/test split. Additional resample IDs can be run when repeated evaluation is +required. +## Installation -There are alternatives: we could perform stratified resamples or cross validate. +The experiment drivers are an optional dependency because running experiments is not +needed to import the classifiers or read the result tables: +```bash +pip install -e ".[experiments]" +``` + +At present, the released `tsml-eval` package pins an older aeon range than this project +uses. If pip reports an aeon dependency conflict, install a checkout of the `main` +branch of `tsml-eval` until a compatible release is available: + +```bash +git clone https://github.com/time-series-machine-learning/tsml-eval.git +pip install -e ./tsml-eval +``` + +Deep-learning classifiers also need the project's deep-learning extra: + +```bash +pip install -e ".[deep-learning]" +``` + +## Run one experiment + +The smallest complete example is +[`run_single_dataset.py`](../multiverse/experiments/run_single_dataset.py). Set the +dataset path and output path, then choose an archive dataset and a classifier name: + +```python +from tsml_eval.experiments import ( + get_classifier_by_name, + load_and_run_classification_experiment, +) + +classifier_name = "ROCKET" +classifier = get_classifier_by_name(classifier_name, random_state=0) + +load_and_run_classification_experiment( + problem_path="C:/Data/Multiverse", + results_path="./results-raw", + dataset="BasicMotions", + classifier=classifier, + classifier_name=classifier_name, + resample_id=0, + overwrite=False, +) +``` + +`problem_path` must contain the standard archive layout: + +```text +//_TRAIN.ts +//_TEST.ts +``` + +The function loads those files, fits only on `_TRAIN.ts`, predicts `_TEST.ts`, and +writes: + +```text +//Predictions//testResample.csv +``` + +With `overwrite=False`, an existing result is retained, which makes interrupted batch +runs safe to restart. The same call can be put inside loops over classifiers, datasets, +and resample IDs; see `run_benchmark.py` for that pattern. + +## Result-file format + +The output is a tsml-format classification CSV, rather than a normal rectangular table. +It is intended to be read with `load_classifier_results`, not parsed with +`pandas.read_csv`. + +The file contains: + +1. A metadata line containing the dataset, classifier, split (`TEST`), resample ID, + time unit, and a description. +2. A parameter-information line containing the estimator configuration. +3. A summary line containing accuracy, fit time, predict time, benchmark time, memory + usage, number of classes, and optional train-error-estimation fields. +4. One line per test case containing the true class, predicted class, one probability + for each class, prediction time, and an optional case description. + +For example, a result for `ROCKET` on `BasicMotions` at resample `0` is found at: + +```text +results-raw/ROCKET/Predictions/BasicMotions/testResample0.csv +``` + +Load and score one file as follows: + +```python +from tsml_eval.evaluation.storage import load_classifier_results + +result = load_classifier_results( + "results-raw/ROCKET/Predictions/BasicMotions/testResample0.csv" +) +result.calculate_statistics() +print(result.accuracy) +print(result.balanced_accuracy) +print(result.fit_time, result.predict_time) +``` + +The stored probabilities allow metrics such as AUROC and log loss +to be calculated later. + +## Collate and compare results with tsml-eval + +Once one result file exists for every requested classifier, dataset, and resample, use +`evaluate_classifiers_by_problem` to collate and compare them: + +```python +from tsml_eval.evaluation import evaluate_classifiers_by_problem + +evaluate_classifiers_by_problem( + load_path="./results-raw", + classifier_names=["ROCKET", "DrCIF", "ConvTran"], + dataset_names=["BasicMotions", "ItalyPowerDemand", "Trace"], + save_path="./evaluations", + resamples=1, # evaluates resample 0 + eval_name="example", + continue_on_missing=False, +) +``` + +The evaluator finds files using the standard directory layout and writes an evaluation +directory containing per-metric CSV files, summary CSV files with mean scores and mean +ranks, and comparison figures. Set `resamples=30` to evaluate IDs `0` through `29`, or +pass an explicit list such as `[0, 1, 2]`. By default a missing file is an error; use +`continue_on_missing=True` when deliberately allowing incomplete comparisons. The +default behaviour removes incomplete datasets from summary comparisons, which keeps +all classifiers on the same set of completed problems. + +## Collate into this repository's result tables + +The repository's `ingest.py` converts the raw tsml-eval files into the smaller tables +used by the leaderboard. It loads each file with `load_classifier_results`, calculates +the standard metrics, and writes one file per classifier and metric: + +```text +results/multiverse//_.csv +``` + +For example: + +```python +from multiverse.experiments.ingest import ingest + +ingest( + classifier="ROCKET", + predictions_path="./results-raw", + datasets=["BasicMotions", "ItalyPowerDemand", "Trace"], + resample=0, +) +``` + +The generated metric files have one row per dataset and one column per resample. Their +index is labelled `Resamples:`; when there is one resample, the single column is usually +`0`. Missing prediction files are left out rather than filled with a score, so failures +remain visible and can be reported separately. + +To ingest the default Multiverse-core dataset list and the configured classifiers, edit +the settings at the top of `multiverse/experiments/ingest.py` and run: + +```bash +python -m multiverse.experiments.ingest +``` + +Finally, build the HTML leaderboard from the collated tables: + +```bash +python -m multiverse.experiments.tables +``` + +The lower-level table API is also available for inspection: + +```python +from multiverse.experiments.tables import load_metric, leaderboard + +accuracy = load_metric("ROCKET", "accuracy") +leaderboard( + datasets=["BasicMotions", "ItalyPowerDemand", "Trace"], + estimators=["ROCKET", "DrCIF", "ConvTran"], + metrics=["accuracy", "balacc", "logloss"], + output_path="./results/multiverse/leaderboard.html", +) +``` + +The ingested tables are a convenient +summary for this repository's leaderboards and should be regenerated if raw results are +changed. diff --git a/docs/leaderboard.md b/docs/leaderboard.md index 2f0c747..c6d673d 100644 --- a/docs/leaderboard.md +++ b/docs/leaderboard.md @@ -1,8 +1,8 @@ # Leaderboards -The leaderboards can be interactively generated on the WEBSITE. These are some +The leaderboards can be interactively generated on the WEBSITE COMING SOON. These are some illustrative static leaderboards ranked on classification accuracy. We will embed -the interactive version and update this dynamic in time. +the interactive version when its ready. In the interim, we present some generative tools. ## Generating a leaderboard @@ -29,9 +29,7 @@ are used, so each column describes the same problems; anything left out is liste the page with the reason taken from `results/multiverse/missing_results.csv`. A critical difference diagram can be added with `critical_difference=True`. It is off -by default because on the current results the omnibus Friedman test does not reject -over the leading estimators, so the diagram is a single clique and shows nothing the -table does not. +by default because we generate the front page table for all estimators. Building the same page from the command line: @@ -57,41 +55,3 @@ Every column in the generated table is sortable: click a heading to sort by it, click again to reverse. The first click puts the best value on top, so ascending for ranks and for log loss, descending for the rest. -**A warning on `max_cd_estimators`.** Truncating the critical difference diagram to the -best `n` estimators changes the statistics rather than just hiding rows. Ranks, the -omnibus test and the corrected alpha are all computed over the subset shown. The -diagram starts with an omnibus Friedman test, and dropping the weakest estimators -compresses the spread of average ranks, which can take that test from rejecting to not -rejecting. When Friedman does not reject, aeon places every estimator in a single clique -and runs no pairwise tests at all, so no differences appear. - -On the current Multiverse-core results this is not hypothetical: - -| Estimators in the diagram | Friedman p | Outcome | -|---|---|---| -| top 6 | 0.44 | one clique, no pairwise tests | -| top 8 | 0.14 | one clique, no pairwise tests | -| all 10 | 0.0003 | 7 significant pairs at alpha/(k-1) = 0.011 | - -Treat a truncated diagram as a statement about that subset only. - -`available_estimators()` lists the estimators that have results, and `load_metric()` -returns one estimator's scores for one metric as a `pandas.Series` if you would rather -build your own table. - -## Multiverse - -## Multiverse-core - - -## EEG archive - -The EEG archive is a collection of EEG classification problems, described in [1]. On -release, it contains 30 datasets. Two of these are univariate and two are not -available on zenodo. The resulting list is contained in the multiverse - - -## UEA archive - -People will still use the UEA archive, so it is worth maintaining a list for sanity -checks. The archive contains 30 datasets, but \ No newline at end of file diff --git a/docs/memory.md b/docs/memory.md index 00e2070..b244c00 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -18,11 +18,8 @@ in for one. ## What is recorded `tsml-eval` writes one `memory_usage` value per classifier, dataset and resample: the -peak memory observed during `fit`. It is in the raw prediction files, and this -repository's ingest brings across only the accuracy-style measures, so it has not been -carried over. - -One number, host side, fit only. +peak memory observed during `fit`. It is in the raw prediction files, but it is only a +proxy for the memory footprint of a classifier. ## Why the figures we hold cannot stand in diff --git a/docs/results.md b/docs/results.md index 2bac006..1eb47f4 100644 --- a/docs/results.md +++ b/docs/results.md @@ -1,56 +1,47 @@ -# Classifier Results +# Classifier results -Results used in past bake offs are available on [tsc.com] -(https://timeseriesclassification.com) and -obtainable in code with [aeon](https://github.com/aeon-toolkit/aeon/blob/main/aeon/benchmarking/results_loaders.py). +This page describes the prediction and summary results stored in this repository. For +the archive collections, dataset splits and dataset selection, see +[`datasets.md`](datasets.md). For running experiments and converting raw predictions +into these tables, see [`evaluation.md`](evaluation.md). -```python - -from aeon.benchmarking.results_loaders import get_available_estimators -cls = get_available_estimators("Classification") # doctest: +SKIP -from aeon.benchmarking.results_loaders import get_estimator_results -cls = ["HC2"] # doctest: +SKIP -data = ["Chinatown", "Adiac"] # doctest: +SKIP -get_estimator_results(estimators=cls, datasets=data) # doctest: +SKIP -``` - -We currently store the multiverse results in the results directory. Currently -only have accuracy for the default splits for subsets of the multiverse. This is -still a work in progress. You will soon be able to explore and download these results -interactively on the [multiverse website] (COMING SOON). +## Published and generated results -The dataset lists are +Results from earlier bake-offs are available from +[timeseriesclassification.com](https://timeseriesclassification.com) and can be +loaded through aeon's [results loaders](https://github.com/aeon-toolkit/aeon/blob/main/aeon/benchmarking/results_loaders.py): ```python - -from aeon.datasets.tsc_datasets import multiverse_core, multiverse2026, eeg2026 -print(len(multiverse_core)) # 66 -print(len(multiverse2026)) # 133 -print(len(eeg2026)) # 28 - +from aeon.benchmarking.results_loaders import ( + get_available_estimators, + get_estimator_results, +) + +estimators = get_available_estimators("Classification") +results = get_estimator_results( + estimators=["HC2"], + datasets=["Chinatown", "Adiac"], +) ``` -### The Full Multiverse, 2026 +The regenerated multiverse results distributed with this repository are stored under +[`published_results/`](../published_results/). Raw prediction files and the scripts +that produce summary tables are described in [`evaluation.md`](evaluation.md). -The full multiverse has 133 datasets in it. We have results for 17 classifiers on -some subset of these problems. +## Repository summary tables -```python -from pathlib import Path -import pandas as pd +The checked-in benchmark tables are arranged by collection, estimator and metric. For +example: -# Run this from the repository root -df = pd.read_csv(Path("results") / "multiverse" / "accuracy_mean.csv") -print(df.head()) +```text +results/multiverse//_.csv ``` -## The Multiverse-core (M-core) -We specify a subset of 66 datasets for evaluation. These are more balanced in -application, remove overly similar, too simple or zero information datasets and -have a good distribution in size and length. +Each table has one row per dataset and one column per resample. The index is labelled +`Resamples:`. A table named `accuracy_mean.csv` may also be present at the collection +level for a compact estimator-by-dataset view. -Per-estimator results for Multiverse-core are stored one directory per estimator, -with one file per performance measure. +Load a metric for one estimator with the repository helpers: ```python from multiverse.experiments.tables import available_estimators, load_metric @@ -60,18 +51,30 @@ accuracy = load_metric("RIST", "accuracy") print(accuracy.shape) ``` -See [`docs/leaderboard.md`](leaderboard.md) for building a ranked leaderboard from -these. +`load_metric` averages across resample columns when more than one resample is present. +Datasets without a result for an estimator are left missing rather than assigned a +placeholder score. The leaderboard uses the common set of completed datasets when +comparing estimators, so the comparison is made on the same problems. +## Building summaries -## The EEG Classification archive, 2026 +After raw tsml-eval files have been ingested, build the HTML leaderboard with: -The EEG archive is a sub-project meant to benchmark EEG classification algorithms. -The project is based around [aeon-neuro](https://github.com/aeon-toolkit/aeon-neuro) +```bash +python -m multiverse.experiments.tables +``` +Or generate one explicitly: ```python -df = pd.read_csv(Path("results") / "eeg" / "accuracy_mean.csv") -print(df.shape) +from multiverse.experiments.tables import leaderboard + +leaderboard( + datasets=["BasicMotions", "ItalyPowerDemand", "Trace"], + estimators=["ROCKET", "DrCIF", "ConvTran"], + metrics=["accuracy", "balacc", "logloss"], + output_path="./results/multiverse/leaderboard.html", +) ``` +See [`leaderboard.md`](leaderboard.md) for the available views and ranking options. diff --git a/docs/runtime.md b/docs/runtime.md index c259acd..6a6071b 100644 --- a/docs/runtime.md +++ b/docs/runtime.md @@ -1,54 +1,26 @@ # Runtime **Coming soon.** This page will hold runtime comparisons across the Multiverse -estimators. Memory is a separate page, [memory](memory.md), for the same reasons and a -few of its own. +estimators. -**We have not yet structured an experiment to compare runtime.** Every run behind the -results in this repository was set up to measure predictive performance. Which partition -a classifier was queued on, how many cores it was given, how many epochs it trained for, -whether a job was retried at a higher memory ceiling: all of those were chosen to get -accurate results out at a reasonable cost, and none were held constant across estimators -because nothing depended on it. The timings that came out are a by-product of that, not -a measurement anyone designed. - -So nothing is published here yet, deliberately. Comparing runtime needs its own -experiment, with the conditions below fixed in advance, and we have not run one. The -rest of this page records what those conditions are, and why the figures we already hold -cannot stand in for them. - -## The measurements exist - -Every run already records timings. `tsml-eval` writes, per classifier, dataset and -resample: +`tsml-eval` writes, per classifier, dataset and resample: - `fit_time` and `predict_time`, in seconds; - `benchmark_time`, the time that machine took to sort 1,000 seeded random arrays of 20,000 elements; -- `memory_usage`, the peak memory during `fit`, which [memory](memory.md) covers. - -They are in the raw prediction files. What this repository ingests under `results/` is -only the accuracy-style measures, one file per metric, so the timings have not been -brought across yet. That is a small piece of work; the reason it has not been done is -below, not the effort. -## Why the figures we hold cannot stand in - -Each of these is a condition a timing experiment would have to fix, and that these runs -left free. +They are in the raw prediction files. **The runs are spread across different hardware.** Multiverse results have been produced -on GPU partitions with H200 and A100 cards and on CPU-only nodes, with different core -counts. GPU jobs in our configurations are allocated two CPUs each. A fit time from one +on Southampton's IRIDIS GPU partitions with H200 and A100 cards. GPU jobs in our configurations are allocated two CPUs each. A fit time from one partition and a fit time from another are two different measurements that happen to share -a unit. +a unit. CPU experiments were run on UEA's HALI cluster. **GPU and CPU methods are not on one axis.** For the deep learners nearly all the work is on the accelerator and the host CPU mostly feeds batches; for the classical ensembles there is no accelerator at all and the time scales with the cores allocated. Comparing them measures the hardware at least as much as the algorithm, and the ratio moves when -either side changes. A statement like "X is 40 times faster than Y" is, in this setting, -a statement about a purchasing decision. +either side changes. **Wall-clock contains things that are not the algorithm.** Queueing, data loading, and retries: our controllers escalate a job's memory request from 64 GB to 128 GB after a @@ -60,28 +32,6 @@ faithful ports of the same paper can differ several-fold on time because the aut picked 500 epochs and the toolkit's default is 2000. Early stopping and best-epoch selection move it again. None of that is a fact about the architecture. -**Which device a run actually used is not reliably recorded.** In the version of the -experiment tooling used for these runs, the device description inspects TensorFlow only, -so a PyTorch estimator reports CPU whether or not it ran on a GPU. Any timing table built -from those records has to have its device column reconstructed from the job -configuration rather than trusted as written. - -## What a fair comparison would need - -- The compared estimators run on the same hardware, or CPU timings normalised by - `benchmark_time`, which exists for exactly this purpose. There is no equivalent - normaliser for GPU work. -- Thread and core counts fixed and recorded, since the classical methods scale with them. -- `fit` and `predict` reported separately. They answer different questions: fit time is - the cost of research, predict time is the cost of deployment, and the ranking is not - the same on both. -- Repeated runs. Timings vary far more between repeats than accuracy does, especially on - shared nodes. -- Asymptotic complexity in the number of cases, series length and channels reported - beside the measured times, so a reader can tell whether a result will hold at a - different scale. -- The device stated per run, from the job configuration. - -Until most of that is in place, this page stays empty. For the same reason the -[leaderboard](leaderboard.md) carries no fit time or predict time columns, and -[memory](memory.md) is empty too. +We can give a crude comparison with the results we have for the CPU and GPU groups independently. However, +an experiment to characterise the run time complexity and a function to estimate expected run time for a specific +configuration in controlled environments would be more useful.