Skip to content

Open question: should FDS support grouping shots (e.g. into sessions)? #87

Description

@NathanCummings

Open question, no proposed answer

Should FDS be able to name and address a set of shots? Filing this to hold the question, not to argue a position — the motivating case is one facility's operating practice and may not generalise.

What prompted it

MAST shots carry three free-text scientific-metadata fields that turn out to describe a structure rather than classify a shot. Measured over the 11,573 shots on the live catalogue:

name distinct values perfectly contiguous in time median shots per value
objective 941 98% 12
program 781 89% 12
experiment_tags 267 14% 10

objective names a session: a consecutive run of about a dozen shots sharing a purpose. 99% of objectives fall under exactly one program, and programmes are research themes:

program:  "ELM control"
  objective: "Determine the effect of the ELM coils in Configuration 8 (Odd parity...)"
  objective: "investigate effect of ELM coil config 9 (Even parity 60 deg phase)"
  objective: "Continue on testing the effect of ELM coils on H-mode..."

experiment_tags behaves differently — scattered rather than contiguous — and is a genuine cross-cutting classifier. The contrast is what makes the first two look structural.

The gap

There is currently no way to group Shots. A Collection groups Datasets (CollectionDataset) and child Collections (CollectionMember), and it belongs to a Shot through shot_id. Nothing in the model holds a set of shots.

So session identity is denormalised: the same objective string is repeated on each of the session's dozen shots, because there is nowhere else to put it.

ADR-0051 decided that a discoverable run is a Collection, but that was settled for an analysis run producing datasets. An experimental session producing shots is a different shape, and whether the precedent extends or breaks is exactly what is undecided.

What is not being claimed

  • That FDS should model campaign → program → objective. That is how MAST ran its operations, not a fact about tokamaks. Nothing device-specific should enter the model on this evidence.
  • That the current handling is broken. Open-vocabulary metadata (ADR-0030) records what the provider asserts, and that is what it is doing. These names are searchable through the ordinary annotation filter today.
  • That contiguity defines a session. It is measurable and it is evidence, not a definition — a detector installed for six months would look contiguous and is not a session.

Things to weigh, whenever this is picked up

  • The capability question ("can a set of shots be named and addressed?") is separable from the semantics question ("what is a session, and does every facility have one?"). The first might generalise; the second probably does not.
  • A session is a plausible citation target, which pulls on the persistent-identifier and PROV decisions.
  • If sessions became first-class, the ingest would create one per objective rather than stamping the string onto every shot — so this also bears on what facility ingestion is expected to produce.
  • Doing nothing is a legitimate outcome. The cost is that a session is reachable only by filtering on a 941-value free-text vocabulary.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    architectureCross-cutting designmetadataCatalogue metadata and semanticsquestionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions