Make sense of animal neuroscience data you did not organise yourself.
A drive arrives. The person who filled it has graduated. Somewhere in it are
three years of recordings, and nobody left can say which animal is which, what
was injected when, or which of four folders called final is the one the
figures came from.
NDOS reads that drive without changing it, and tells you what is on it — what it observed, what it inferred, and what it could not determine. It reconstructs the experiment behind the files: animals, surgeries, sessions, acquisitions and the analyses built on them. Then it lets you ask scientific questions of the result and see the evidence behind every answer.
Four things it does that a file browser cannot:
- Recovery. Inventories heterogeneous storage, reads inside archives without extracting them, and proposes a structure — showing the rule and the evidence behind every guess, so you can correct it rather than trust it.
- Search. Finds the word you remember —
CA1,GCaMP, a construct name — across filenames, lab notes, protocol files and spreadsheets, including.docxand.xlsx. A hit in a surgery log reports which animals it names, which is what turns a document into a route to the recordings. - Discovery. Builds cohorts from incomplete records and returns three answers, not two: matched, excluded, and cannot be ruled out. A session whose species nobody wrote down is not a session known not to be a mouse, and a tool that conflates those hands you a biased cohort without saying so.
- Traceability. Records what produced a result, and traces a figure back through the commands and parameters to the raw files it came from.
Nothing changes what is already on your disk unless you name a command that
does. Three can, and each shows a plan and waits: organize --mode move,
tags sweep --apply, protect --apply. Other commands write new files — a
manifest, metadata sheets, a search index, a layout of symlinks, a BIDS export
— always at a path you give them, never over your originals.
The layout is one part of this, not the whole of it.
SPECIFICATION.md
states a directory structure, session shape, identifiers and data flags that a
project can be checked against, and ndos validate reports whether one
conforms. It is a practical convention rather than a rival to BIDS or NWB —
NDOS hands off to both — and it is deliberately separable from everything
above, so a lab that has already chosen a layout can still use the rest.
Status: 0.1, and honest about it. Everything documented below works and is tested against real lab storage on Linux, macOS and Windows. The specification is a draft. Interfaces may change. No external lab has completed a pilot yet — that is what we are looking for.
Four sessions, four animals, one question — and three different answers. This output is committed in the repository, so it is what the code actually prints:
Considered 4 sessions: 1 matched, 2 unresolved, 1 excluded.
MATCHED (1)
ndos-0000000001 raw_data/M101/20250310
species = mus musculus (declared)
sex = F (declared)
target_region = dorsal CA1 (declared)
CANNOT BE RULED OUT (2)
These meet every criterion that could be checked, but a deciding
value was never recorded. They are not non-matches.
ndos-0000000003 raw_data/M103/20250310
sex: recorded as unknown; checked but could not be determined
ndos-0000000004 raw_data/M104/20250311
species: never entered
EXCLUDED (1)
1 excluded by species=mouse
M102 is a rat — recorded evidence contradicts the query. M103 and M104
are the reason this exists. A tool reporting only matched and excluded would
say 1 of 4 and look clean; the truth is one match and two sessions nobody
can decide, for two different reasons — one was checked and came back
unknown, the other was never filled in.
Treating either as "not a mouse" biases the cohort towards whichever animals happened to have fuller records, and says nothing about having done so.
Why each animal lands where it does is spelled out in examples/cohort-demo/. A test regenerates that output and fails if it drifts, so what you just read is what the code does.
To run it yourself — the example ships in the repository, so clone rather than install:
git clone https://github.com/Elnazkarami/N-DOS.git ndos && cd ndos
python3 ndos.py table check examples/cohort-demo/metadata --emit linked.json --include-empty
python3 ndos.py query linked.json -w species=mouse -w sex=F -w 'target_region~CA1'If you have a directory of lab data nobody fully understands any more, that is exactly what this needs to meet. It takes about fifteen minutes and will not move or change your data.
| Tell us how it went → | Even "I gave up at step three" is a result |
| It read my folders wrongly → | The most useful report there is |
| Something is broken → | A crash or a wrong answer |
No data needs sharing — every form asks about the shape of your directories, never their contents. Start at QUICKSTART.md.
NDOS Core is Python 3.9+ standard library only. No pip install, no
environment, no dependencies. This is deliberate: the machine that most needs
to be inventoried is often an acquisition PC where you are not allowed to
install anything.
git clone https://github.com/Elnazkarami/N-DOS.git ndos
cd ndos
python3 ndos.py --help
python3 ndos.py report /path/to/your/dataEvery module is also a standalone script — python3 ndos_report.py ... works
identically. They are not single files you can lift out individually, though:
most import their siblings, so run them from the checkout.
If you would rather type ndos report than python3 ndos.py report:
pip install ndos
pip install -e . # or from a cloneEither adds the command and nothing else — there are no dependencies to install, and CI fails the build if that ever stops being true. New here? Start with QUICKSTART.md. When NDOS reads your data wrongly — and on some layout it will — RECIPES.md is what to do about it.
ndos_scan.py and ndos_report.py never modify, move, rename, extract, or
delete anything below the directory you point them at. They open files only
to read bytes for checksums. This is verified by tests that snapshot every
size and modification time before and after a scan.
Nothing in NDOS Core writes to your data without an explicit, reviewable plan that you approve first.
Starting from a directory nobody understands:
python3 ndos.py report /path/to/chaos # what is in here?
python3 ndos.py archive inspect /path/to/chaos -c arch.json # what is in the zips?
python3 ndos.py search index /path/to/chaos -i search.db # make it searchable
python3 ndos.py search find "CA1" -i search.db # where is that word?
python3 ndos.py organize apply /path/to/chaos -d ./project # build the N-DOS layout
python3 ndos.py table export ./project -d ./metadata # fill in what only you know
python3 ndos.py table check ./metadata --emit linked.json
python3 ndos.py query linked.json -w species=mouse -w target_region=CA1
python3 ndos.py convert bids ./project -d ./bids-export --writeNothing in that sequence changes your data. It writes a manifest, a search
index, metadata sheets and an export, each at a path you named, and the layout
is built from symbolic links pointing at your originals. The commands that can
change what is already on disk are organize --mode move, tags sweep --apply and protect --apply; none of them is above, and each shows a plan
and asks first. (archive extract writes new files too, but never alters the
archive it read.)
The order below is the order you meet them: understand what you have, then structure it, then describe it, then use it.
Point it at a folder and get a readable account of what is in there, what is duplicated, how it appears to be organised, and what needs attention. Requires no metadata and no prior setup.
python3 ndos_report.py /path/to/data # readable summary
python3 ndos_report.py /path/to/data -f markdown -o report.md
python3 ndos_report.py manifest.json -f json # machine-readableIt reports:
| Section | What it answers |
|---|---|
| What is in here | Composition by category and size — ephys, imaging, behaviour, analysis, archives |
| Inferred folder structure | Which directory level looks like a subject, a session, a date — and which names contradict that |
| Duplicate files | Byte-identical copies and how much space they waste |
| Largest directories | Where the volume actually lives |
| Needs attention | Unextracted archives, zero-byte files, unreadable paths, names that break on other systems |
Compound formats are recognised where a bare extension is not enough:
M01_g0_t0.imec0.ap.bin is reported as electrophysiology, not as an anonymous
.bin. Genuinely ambiguous extensions are labelled ambiguous rather than
guessed at — a confident wrong label is worse than an honest unknown.
Inferred structure is always labelled as inference. NDOS distinguishes what it observed, what it computed, and what it guessed, and never presents one as another.
Produces a versioned JSON manifest: every file with its path, size, modification time, and SHA-256 digest, plus an explicit list of everything skipped and why.
python3 ndos_scan.py /path/to/data --output manifest.json
python3 ndos_scan.py /path/to/data --no-checksum # faster, no dedup
python3 ndos_scan.py /path/to/data --exclude '*.tmp'The manifest is the input to every other NDOS module, so a slow checksummed
scan only has to happen once. The output contract is versioned in
schemas/manifest.schema.json.
An inventory that quietly omits data is worse than no inventory, so
unreadable directories, permission failures, and symlinks are recorded in a
skipped list rather than silently dropped.
Labs zip their archives because the data is enormous. On a real lab drive,
88% to 100% of everything was inside .zip files, which meant no
inventory could describe any of it.
This reads an archive's index rather than its contents, so a 300 GB collection can be catalogued without unpacking a byte:
python3 ndos_archive.py inspect /Volumes/archive -c archives.json
python3 ndos_archive.py search archives.json '*.avi'Listings are cached and keyed on size and modification time. That matters: on a slow external drive the first read of a 2 GB archive took nearly a minute, and nobody should pay that twice.
Extraction is planned, then confirmed, then done. Ask for what you want, see exactly what it would write and what it would cost, and only then say yes:
python3 ndos_archive.py plan archives.json --dest ./work --name '*.avi'
python3 ndos_archive.py extract archives.json --dest ./work --name '*.avi'The plan reports how many files, how many bytes, which already exist, and whether there is enough free space — refusing to start if there is not. Nothing is written until you confirm.
Members that would escape the destination are refused. An archive can
contain ../../etc/something, and unpacking one you did not create is a real
way to get files written where you did not intend. Those members are listed
under REFUSED and never extracted.
.tar and .tar.gz need --include-tar, because unlike a zip they must be
streamed end to end to be listed, which on slow storage is expensive enough to
be a deliberate choice.
An inventory describes a directory. It does not make an inherited archive understandable to a PI whose data collector left years ago. This builds the missing half: the N-DOS project, derived from the original structure rather than imposed on it.
python3 ndos_organize.py plan /path/to/chaos -d ./project # see it first
python3 ndos_organize.py apply /path/to/chaos -d ./project # then confirm
python3 ndos_organize.py undo ./project/.ndos-layout-log.jsonThe layout is the one defined in the N-DOS manuscript:
project/
├── raw_data/<SubjectID>/<SessionID>/ # acquisition files
├── processed_data/ analysis/ derivatives/
├── flagged_data/ figures/ scripts/ metadata/
└── README.md
SessionID follows the manuscript's YYYYMMDD convention, becoming
YYYYMMDD_01, _02 when a subject was recorded more than once that day.
Nothing is copied or moved by default. The tree is built from symbolic
links, so a 45 GB collection is organised in seconds, occupies about 40 KB,
and is undone by deleting it. --mode copy and --mode move exist, are
planned and confirmed the same way, and move says plainly that it relocates
your data. Every apply writes .ndos-layout-log.json, and undo reverses it —
including moving files back.
Filenames follow the manuscript's conventions, SubjectID_SessionID_type:
A0634_20201122_video-0.avi A0634_20201122_raw-info.rhd
A0634_20201122_position-Take-2020-11-22-06.32.30-PM.csv
A0634_20201122_experimenter-notes.csv
The data type comes from the file itself — its extension, then its name — and
never from the directories above it, which decide the role instead. Where no
standard type applies, the original descriptor is kept (_analogin.dat)
rather than forcing a file into a category it may not belong to: a confident
wrong label is worse than an unfamiliar one, because the filename is what
everyone reads first. Pass --keep-original-names to skip renaming entirely.
Because the default mode is links, renaming costs nothing and risks nothing: the link carries the standard name while the file it points at keeps its own.
Sessions are clustered, not split. A miniscope starting at 18:32:25 and an
Intan at 18:32:20 are one recording, not two. Acquisition times within ten
minutes are treated as one session; genuinely separate recordings on a day
become YYYYMMDD_01 and _02.
Every placement explains itself:
raw_data/A0634/20201122/0.avi
from 0.avi
why folder 'A0634' matches a subject identifier;
'A0600' above it read as a cohort or range
why date '2020_11_22' with acquisition time '18_32_25'
why imaging and video is treated as raw_data
Structure is read from directory names, from filenames when the folders are
silent (2020_11_20.zip carries its date nowhere else), and from compound
names like A0634_201122_183220.
Where the folders record a session but not when it happened, --dates takes
the date from your sessions.csv, turning ses-01 into 20250314. A dated
identifier is what the standard establishes, because it places a recording in
time without opening anything. A date already in the path is never
overridden.
Analysis-tool output is left alone. A Phy or Kilosort sorting is opened by
looking for spike_times.npy and params.py by name, so renaming inside one
would stop the tool reading it back. Directories written by Phy, Kilosort,
SpikeInterface, suite2p, Open Ephys, DeepLabCut or Zarr are recognised by the
files those tools require, placed under the session they belong to, and carried
across with their filenames and internal structure untouched. The plan says
when it has done this and why.
Nothing is dropped. Files whose subject or session cannot be determined go
to flagged_data/ with their original structure intact and a
flagged_notes.json saying why — which is what the N-DOS layout reserves that
directory for.
Redundant copies are recognised, not duplicated. A lab that restructured its data once has the same recording in two places; on a real drive this was 3.0 GB. Each file is linked once and the copies are listed, rather than appearing under invented names as though they were distinct.
Scanning reveals what is on disk. It cannot reveal which animal a recording came from, what was injected, or when. That knowledge lives in a notebook or an Excel sheet, so NDOS meets it there.
python3 ndos_table.py export manifest.json -d metadata/ # build the sheets
# ... open them in Excel and fill in the blanks ...
python3 ndos_table.py check metadata/ --emit linked.jsonMetadata lives in three linked tables, because a fact recorded once should govern every session it applies to:
| File | One row per | Filled by |
|---|---|---|
animals.csv |
animal — species, strain, sex, date of birth, genotype | you, once per animal |
procedures.csv |
surgery, injection, implant, drug, training | you; NDOS cannot observe a surgery, so it never touches this file |
sessions.csv |
recording session — date, task, QC | NDOS pre-fills what it observed; you add the rest |
animals.csv is seeded with the subject names NDOS found in your folder tree,
so you start with rows rather than a blank sheet. Sessions are regenerated on
every export with observed columns refreshed and typed-in values carried
across; a row that disappears is reported rather than silently taking its
metadata with it.
Intervals are computed, never typed. Because a procedure has a date and a
session has a date, NDOS derives days_since_injection, days_since_implant,
age_days, and so on. These carry the status computed, so they are never
mistaken for something a person asserted — and they make "recorded three to
five weeks after the injection" a query you can actually run.
Entry is forgiving, validation is strict. mouse, Mouse, and mice all
resolve to mus musculus; ephys to electrophysiology; viral injection
to injection; female to F. What you typed is preserved beside the mapped
value so the mapping can be audited. But 21/03/2025 is refused, because
03/04/2025 means 3 April in the UK and 4 March in the US and guessing would
silently corrupt a date.
Validation spans the tables, not just each file: a session naming an animal with no row, or a procedure for a subject nobody described, is reported as a broken link.
A blank cell and the word unknown mean different things, and NDOS keeps them
apart: blank means nobody has filled it in yet, unknown means somebody
checked and could not determine it. --emit writes evidence-typed records
against schemas/session_metadata.schema.json.
The layout reserves flagged_data/ and temp/, and the standard defines the
flags that make them mean something:
{"validated": true, "temp": false, "deletable": false}.
python3 ndos_tags.py set spikes.npy --validated --note "curated in Phy"
python3 ndos_tags.py list ./project --flag temp
python3 ndos_tags.py sweep ./project # plan a cleanup
python3 ndos_tags.py sweep ./project --apply # after reading itTags live in a tags.json beside the data they describe, one per session, so
a session directory stays self-describing if it is moved or copied.
A validated file is never swept, whatever its other flags say, and marking
something both validated and deletable is recorded as a conflict rather than
resolved silently. sweep plans by default and deletes only on --apply with
a confirmation: the manuscript imagines a maintenance script removing
temporaries automatically, and deletion driven by a hand-edited flag is how a
lab loses data it meant to keep.
Files that look like scratch but were never flagged are listed separately
and only with --include-untagged, because nobody has vouched for them.
Scratch is judged relative to the project root, so a project living under
/tmp does not have all of its files called temporary.
ndos_organize tags automatically: spike-sorting scratch is routed to
processed_data/<sub>/<ses>/temp/, flagged temp, and recorded in that
session's derived_metadata.json — so a sweep finds it later without anyone
remembering which files were intermediates.
python3 ndos.py validate ./my-study This project does not yet conform.
REQUIRED (2)
session 'March14' is not YYYYMMDD or YYYYMMDD_NN
at raw_data/M123/March14
fix rename it, or rebuild the layout with ndos organize
Requirements and recommendations are kept apart. A requirement unmet means the project does not conform and the command exits non-zero, so it can gate a hand-off or a submission. A recommendation — raw data still writable, files not following the naming convention, no metadata yet — is reported and never affects the exit code, because a lab may depart from those deliberately.
Every finding says what to run to fix it.
The standard asks that raw data be set read-only once acquired. A recording is the one thing in a project that cannot be regenerated, and it is usually lost not to a disk failure but to a script writing where it meant to read.
python3 ndos.py protect ./my-study # what it would change
python3 ndos.py protect ./my-study --apply
python3 ndos.py protect ./my-study --check # has anything become writable?
python3 ndos.py protect ./my-study --releaseOnly raw_data/ by default; processed_data/ is meant to change. Permissions
are changed and contents never are. --check exits non-zero when something
that should be read-only is not, so a scheduled job or a pre-publication check
can use it.
Releasing is as easy as protecting, deliberately: a protection people cannot undo is one they work around by copying data somewhere unprotected.
python3 ndos_query.py metadata.json -w species=mouse -w sex=F \
-w 'target_region=CA1' -w 'session_date>=2025-03-01' \
-w 'modalities~electrophysiology' \
--save-cohort cohort.json --name ca1-ephys-spring-2025Operators are = != > < >= <= and ~ (contains). field=* requires
any value; field=? finds values recorded as unknown. Queries are normalised
the same way the data was, so species=mouse finds sessions recorded as
mus musculus, and the report shows you that substitution.
Results come in three groups, not two. A session whose sex was never recorded is not a non-match — it is an open question, and quietly dropping it biases the cohort in a way nobody sees. So NDOS reports:
| Group | Meaning |
|---|---|
| Matched | Every criterion satisfied, each citing the field, value, and evidence status it rests on |
| Cannot be ruled out | Met everything checkable, but a deciding value was never recorded |
| Excluded | Ruled out by evidence that is recorded |
Unresolved sessions stay out of a saved cohort unless you pass
--include-unresolved, and are labelled if you do.
An empty result explains itself rather than returning nothing. It names the constraint that eliminated the most sessions, and distinguishes "your query was too narrow" from "this column was never filled in" — which need opposite fixes.
Cohorts are frozen with the full query plan against
schemas/cohort.schema.json, including counts of
what was excluded and what could not be decided, so a selection can be re-run,
audited, or disputed later.
ndos query needs you to know which field holds the answer. On an inherited
drive you often do not. What you remember is a word — CA1, GCaMP, the name
of a construct — and the thing that knows where it applies is a surgery log in
a spreadsheet nobody has opened in three years.
NDOS already counted that spreadsheet. This reads it.
ndos search index /path/to/drive # read-only, writes an index elsewhere
ndos search find "CA1 AND injection"Results come back in two parts, because a file whose name matches and a document that discusses the word are different claims:
FILES WHOSE NAME OR PATH MATCHES — 14 file(s)
2025-03-14/M01/ses-01/raw/ 5 file(s)
e.g. M01_ses01_g0_t0.imec0.ap.bin, …ap.meta, …lf.bin
subject M01 · session 20250314
backup/2025-03-14/M01/ses-01/raw/ 4 file(s)
e.g. M01_ses01_g0_t0.imec0.ap.bin, …ap.meta, behaviour.csv
subject M01 · session 20250314
DOCUMENTS AND RECORDS MENTIONING IT — 2
surgery_log.xlsx (651 B)
subject_id procedure target construct M123 injection [CA1] AAV9-GCaMP6f …
names M123, M124
observed, from document
procedures.csv (P01)
procedure_id P01 subject_id M123 procedure_type injection target_region dorsal [CA1] …
subject M123
declared, from metadata
Every file is findable by name, whatever is inside it — which matters
because most of a real drive is .bin, .avi, .tif and .dat, and that is
the data. Name matches are summarised by directory rather than listed: a folder
holding 240 matching files should say so, and seeing backup/ appear beside
the original is usually the point.
Then the document half follows the chain the rest of the way. A hit reports which animals the text names, and then where each of those animals' recordings actually are:
surgery_log.xlsx (651 B)
subject_id procedure target construct M123 injection [CA1] AAV9-GCaMP6f M124 injection CA3 …
names M123, M124
observed, from document
→ M123 has 2 sessions:
raw_data/M123/20250314 (2025-03-14, 12 files, electrophysiology, qc pass)
raw_data/M123/20250321 (2025-03-21, 1 file, electrophysiology, qc fail)
→ M124 has no sessions recorded
That last line is often the one that matters: the log says M124 was injected, and nothing on the drive is filed under it. NDOS already links animals to procedures to sessions, so the join existed — this follows it.
Each result also says how it is known: observed means the text is in a file
on disk, declared means a person entered it in a metadata table.
It reads .txt, .md, .json, .yaml, .csv, .tsv — and .docx and
.xlsx, which are ZIPs of XML and so readable without a dependency. .pdf
is not read, because that would mean bundling a parser, and the output says so
rather than returning a quiet empty result.
Flags are searchable, and shown. A file ndos tags marked validated,
temp or deletable carries that into search, along with the note whoever
set it wrote — which is the only place a person explains a judgement. So
ndos search find deletable lists what is queued for removal, and
ndos search find "rig log" finds the recording somebody checked against it.
It also matters in the other direction. Searching CA1 tells you which of
the hits you can trust:
processed_data/M123/20250314/temp/ 1 file(s)
e.g. CA1_draft.npy
subject M123 · session 20250314 · all looks-like-scratch
raw_data/M123/20250314/ 1 file(s)
e.g. M123_20250314_raw.dat
subject M123 · session 20250314 · all validated
looks-like-scratch is NDOS noticing a conventional scratch name, not a flag
anybody set — the two are reported separately and never merged, because one is
an observation and the other is a person's judgement.
AND, OR, NOT, "quoted phrases" and prefix* all work. A search that
finds nothing exactly is retried as a prefix — lab vocabulary is full of
suffixed names like GCaMP6f, which is one token to a search engine — and it
tells you when it did that.
Search ranks; it does not decide. A document mentioning CA1 does not
establish that a session targeted it. To select sessions on recorded evidence —
and to see which ones cannot be ruled out — use ndos query. The output says
this too, at the bottom of every result set.
python3 ndos_query.py linked.json --config analyses.json{"analyses": [
{"name": "theta-power-CA1",
"requires": ["target_region=CA1", "days_since_injection>=21", "qc_status=pass"]}
]}Each analysis is a saved query, so the answer carries the same evidence rules: sessions that qualify, sessions that cannot be ruled out, and why nothing matched when nothing does.
Provenance normally goes uncaptured because capturing it means instrumenting analysis code, and nobody rewrites a working script to satisfy a data policy. So NDOS wraps the command instead:
python3 ndos_prov.py run --input raw/ --output processed/ \
-- python3 preprocess.py raw processedNothing about preprocess.py changes. NDOS checksums the inputs, watches the
output directories before and after, records the git commit and environment,
and writes a run record. The wrapped command's exit code passes through, so
this composes inside existing shell scripts.
Then walk backwards from any result to the data behind it:
python3 ndos_prov.py trace figures/figure1.svgfigure1.svg
└── → figure [run-157f69ae502a] 2026-08-20T02:39:50Z
$ python3 scripts/make_figure.py processed figures
├── M01_ses01_filtered.dat
│ └── → preprocess [run-418c7c94596d] 2026-08-20T02:39:49Z
│ $ python3 scripts/preprocess.py raw processed
│ ├── M01_ses01.dat
│ │ [raw] not produced by any recorded run
Failed runs are recorded too, and marked. "This figure came from a script that exited 1" is exactly what someone needs warning about, so a partial output is captured rather than discarded.
Provenance never collects credentials. Environment variables are recorded
only when named explicitly with --record-env; capturing the whole
environment would routinely bury API keys inside files meant to be shared.
--anonymous additionally omits hostname and username, for provenance you
intend to publish.
Records validate against
schemas/provenance.schema.json and follow
the W3C PROV shape of an activity that used and generated artifacts.
NDOS does not reimplement either standard. NWB conversion is solved by maintained tools, and rewriting it here would produce a worse converter nobody maintains. This prepares the handoff instead:
python3 ndos_convert.py bids ./project -d ./bids-export --write
python3 ndos_convert.py nwb ./project --metadata linked.json --save nwb-plan.jsonbids builds a BIDS-shaped tree of links with dataset_description.json,
participants.tsv and per-file sidecars, mapping N-DOS entities onto BIDS
ones — A0634 → sub-A0634, 20201122 → ses-20201122, and the
discriminator onto acq-. Scratch in temp/ is never published.
It is BIDS-shaped, not validated BIDS, and says so in its own output. Animal electrophysiology is covered by BEP032, which is not finalised, so no export can honestly claim conformance today.
nwb emits a conversion plan with metadata already mapped onto NWB's fields,
ready for NeuroConv or a lab script — and reports what is still missing
(species, sex, date_of_birth) rather than inventing it.
Creates a project profile and the directories NDOS owns. It does not migrate, move, or restructure your data; raw data stays exactly where it is.
python3 ndos_init.py ~/projects/my-studyNot everyone in a lab works at a command line, and the person who knows what the data is often is not the person who is comfortable there.
python3 ndos.py gui # or: ndos guiThat serves a local page and opens it. Four things it does, each calling the same functions the commands call rather than a second implementation that could disagree with them:
| This machine | Pick a folder, see how long a scan would take, watch it run, read the result, check it against the standard |
| Manifest | Open a manifest.json by name and read what a scan found |
| Query | Build a cohort, and see all three answers — matched, excluded, and cannot be ruled out |
| Validate | Check a project's metadata tables against each other, and enter the facts only a person knows |
The query page is the one worth understanding. A session that never recorded a species is not a session known not to be a mouse, and a page that showed you only "matched" and "everything else" would let you publish a cohort that is quietly biased towards the animals somebody happened to write down. So the query runs where the rules live, and answers with all three groups, the constraint that blocked each session, and what filling it in would change.
Nothing the page does writes to your data. Linking a project's metadata for a
query builds the records in memory — ndos table check --emit writes them to
a file, and the page needs no such file to exist.
It is part of the package, already built, so there is still nothing to install: no Node, no npm, no build step, and nothing fetched while it runs.
It is not a website and cannot be made into one. The server binds to the
loopback address, requires a token that only the page it opened is given,
refuses any request that does not claim a local Host, and refuses any request
carrying a foreign Origin — which is what stops a page you happen to have
open in another tab from reading your disk. Nothing is exposed to the network,
and nothing leaves the machine.
The command line remains the primary interface: everything the page can do, a command can do, and some things only a command can do.
A synthetic messy lab project is included, containing problems chosen because they recur in real labs: a duplicated backup, an unextracted archive, a failed acquisition left as a zero-byte file, inconsistent subject naming, and histology stored away from the recordings it validates.
python3 tests/fixtures/make_messy_lab.py /tmp/messy-lab
python3 ndos_report.py /tmp/messy-lab
python3 ndos_table.py export /tmp/messy-lab -d /tmp/metadata
# fill in a few rows across the three sheets, then:
python3 ndos_table.py check /tmp/metadata --emit /tmp/linked.json
python3 ndos_query.py /tmp/linked.json -w species=mouse -w sex=Fpython3 -m unittest discover -s tests -vThe suite runs on the standard library alone. If jsonschema is installed, it
additionally validates generated manifests against the published schema.
| Path | Contents |
|---|---|
ndos.py |
One command dispatching to all of the below |
QUICKSTART.md |
Fifteen minutes, for a lab member |
RECIPES.md |
What to do when a guess is wrong |
ndos_report.py |
Inventory report |
ndos_scan.py |
Read-only inventory |
ndos_archive.py |
Archive inspection and planned extraction |
ndos_organize.py |
Rebuild the N-DOS layout from existing structure |
ndos_table.py |
Linked metadata tables |
ndos_tags.py |
Validation flags, cleanup, validated-file index |
ndos_protect.py |
Make raw data read-only after acquisition |
ndos_query.py |
Cohort queries with evidence citation |
ndos_search.py |
Full-text search over notes, logs and metadata |
ndos_prov.py |
Run provenance and lineage tracing |
ndos_convert.py |
BIDS and NWB handoff |
ndos_init.py |
Start a project, and make a session folder |
ndos_validate.py |
Check a project against the standard |
ndos_gui.py |
Serves the local interface, on this machine only |
ndos_gui_static/ |
That interface, already built, so no Node is needed |
gui/ |
Its source, for changing it — see gui/README.md |
scripts/ |
Regenerating what is derived: field definitions, the built interface |
SPECIFICATION.md |
The standard itself |
schemas/ |
Versioned JSON Schema contracts |
tests/ |
Test suite and synthetic fixtures |
LEGACY.md |
Where the superseded prototypes went, and why |
Scripts that predate NDOS Core are no longer in this branch: they move data with no rollback and one of them contradicts the standard. They remain reachable, with their history — see LEGACY.md.
In scope: non-human animal neuroscience — colony and cohort context, surgeries, injections, implants and drugs, behavioural training, extracellular electrophysiology and calcium imaging, tissue collection and histology, and the analyses derived from them.
Not in scope: replacing BIDS for human MRI/MEG/EEG, replacing NWB as a container, or becoming a public archive. NDOS reads and hands off to those ecosystems rather than competing with them.
Apache-2.0. Chosen over MIT for its explicit patent grant, which institutional legal review generally looks for before adoption.
Earlier revisions of this repository declared MIT in the README. That grant still stands for those revisions; anyone who took the code under MIT keeps it. NDOS Core is Apache-2.0 going forward.