Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions .claude/settings.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,24 @@
{
"extraKnownMarketplaces": {
"quadbio": {
"source": {
"source": "github",
"repo": "quadbio/claude-plugins"
}
}
},
"enabledPlugins": {
"analysis-workflow@quadbio": true
},
"permissions": {
"deny": [
"Bash(git add -A:*)",
"Bash(git add --all:*)",
"Bash(git reset --hard:*)",
"Bash(git clean -f:*)",
"Bash(git clean -fd:*)",
"Bash(git push --force:*)",
"Bash(git commit --no-verify:*)"
]
}
}
4 changes: 0 additions & 4 deletions .gitattributes

This file was deleted.

2 changes: 0 additions & 2 deletions .github/workflows/lint.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -26,5 +26,3 @@ jobs:
frozen: true
- name: Run pre-commit checks
run: pixi run pre-commit run --all-files
- name: Verify notebooks contain no outputs
run: git ls-files -z '*.ipynb' | xargs -0 -r pixi run nbstripout --verify
17 changes: 16 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -150,8 +150,23 @@ dmypy.json
*.gmt
*.gmx

# Directories to ignore
*.zarr

# Data never enters git; only its READMEs and placeholders do.
data/**
!data/**/
!data/**/README.md
!data/**/.gitkeep

# Directories to ignore. Unanchored, so they also match inside every analysis task directory
# (analysis/<topic>/.../<name>_vN/): the untracked half of a task, anchored to the main checkout.
figures/
outputs/
logs/

# A task's results/ (small evidence tables) is tracked even though *.csv / *.txt are ignored above.
!analysis/**/results/**/*.csv
!analysis/**/results/**/*.txt

# OS specifics
**.DS_Store
Expand Down
37 changes: 17 additions & 20 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,39 +1,36 @@
# AGENTS.md — working conventions

This file owns the working conventions and is the canonical guidance for any coding agent.
`README.md` is the user-facing overview; anything documented there is referenced from here,
never restated.
This file is the canonical guidance for any coding agent. `README.md` is the user-facing overview;
anything documented there is referenced from here, never restated.

## Layout
## Workflow

- **Notebooks**: `analysis/[INITIALS]-[YYYY]-[MM]-[DD]_description.ipynb`
- **Data**: `data/<dataset>/{raw,processed,resources,results}/`, gitignored
- **Package**: `src/<package>/`, installed editable from the checkout
The human and agent lanes, task lifecycle, output paths, `data/` layout, working objects and their
write-back are owned by the [analysis-workflow](https://github.com/quadbio/analysis-workflow)
plugin, enabled in `.claude/settings.json`: load its skill before writing analysis code, outputs or
data. Pull requests are reviewed against [`REVIEW_GUIDE.md`](REVIEW_GUIDE.md).

## Paths
## This repo

Never hardcode a path into `data/` or `figures/` — every path hangs off `FilePaths`:
<!-- Replace with this project's facts: its datasets and where each working object lives,
environment specifics, companion code packages. Rules the plugin owns are not restated here. -->

- **Package**: `src/<package>/`, installed editable from the main checkout
- **Paths**: every dataset path hangs off `FilePaths` in `src/<package>/_constants.py`:

```python
from myanalysis import FilePaths

FilePaths.DATA # data/
FilePaths.FIGURES # figures/ — curated output: talk and paper figures
FilePaths.EXAMPLE_DATASET / "processed" / "adata.h5ad"
FilePaths.EXAMPLE_DATASET / "processed" / "adata.zarr"
```

`FilePaths.ROOT` is resolved from git, so it names the *main* checkout even when called from a
worktree and shared data does not follow your branch. Add a dataset as a constant in
`_constants.py`; each one keeps the `{raw,processed,resources,results}` layout by convention.

## Environments

Dependencies live in `pixi.toml`, not `pyproject.toml` — the latter carries package metadata and
the test config. Run `pixi install` after pulling a change to `pixi.toml`, in the main checkout.
Dependencies live in `pixi.toml`, not `pyproject.toml`, which carries package metadata and the
test config.

| Task | Command |
| --- | --- |
| Run Python | `pixi run python script.py` |
| Run tests | `pixi run test` |
| Add conda package | `pixi add <package>` |
| Add PyPI package | `pixi add --pypi <package>` |
| Add a conda / PyPI package | `pixi add <package>` / `pixi add --pypi <package>` |
41 changes: 23 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -83,14 +83,13 @@ Then finish by hand:

```bash
pixi install # create environment from pixi.toml
pixi run install-hooks # pre-commit hooks + notebook output-stripping filter
pixi run install-hooks # pre-commit hooks
pixi run install-kernel # register Jupyter kernel
```

> `install-hooks` also sets up the [nbstripout](https://github.com/kynan/nbstripout)
> git filter. Notebook outputs are then stripped from commits automatically while
> staying in your working copy. **Run it once in every clone** (including remote
> servers and worktrees), or outputs may slip into git.
> **Notebook outputs are committed.** They are the record of what a notebook actually
> produced, and GitHub renders them. Keep them small: clear a notebook by hand before
> committing if it carries a huge embedded image or an accidental dump.

💡 **Tip**: Use `pixi shell` to enter the environment interactively—then you can run commands directly without the `pixi run` prefix.

Expand Down Expand Up @@ -126,6 +125,19 @@ git push

---

## 🤖 Working with coding agents

Agents follow [`AGENTS.md`](AGENTS.md) and the
[analysis-workflow](https://github.com/quadbio/analysis-workflow) Claude Code plugin (a skill plus
guard hooks) that this repo enables. Install the plugin once per machine:

```bash
claude plugin marketplace add quadbio/claude-plugins
claude plugin install analysis-workflow@quadbio
```

---

## ☕ Daily Workflow

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Consider how much of this is needed, given what's in the skill now.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cut to the install step, which is the one thing the skill can't tell a human: it only loads after installation. Dropped the package sentence (pixi.toml shows it) and the REVIEW_GUIDE.md pointer (AGENTS.md has it). a8ce86b


```bash
Expand Down Expand Up @@ -187,12 +199,12 @@ Or edit `pixi.toml` directly and run `pixi install`.
<summary><strong>📓 Data and notebook conventions</strong></summary>

- **Notebook naming**: `[INITIALS]-[YYYY]-[MM]-[DD]_description.ipynb`
- **Data layout** (one folder per dataset):
- `data/<dataset>/raw/` — original data files
- `data/<dataset>/processed/` — preprocessed data
- `data/<dataset>/resources/` — reference data, annotations
- `data/<dataset>/results/` — analysis outputs
- **Figures**: `figures/` or `data/<dataset>/results/`
- **Data layout** (one folder per dataset, gitignored):
- `data/<dataset>/raw/` — bytes as they arrived; never written by analysis
- `data/<dataset>/resources/` — curated inputs that are not data: gene lists, marker tables
- `data/<dataset>/processed/` — objects code loads to do new work, as AnnData `.zarr`
- `data/<dataset>/results/` — a notebook's own outputs, file names prefixed with the notebook's stem
- **Figures**: `figures/<topic>/`
- **Import paths** via the local package:

```python
Expand All @@ -212,13 +224,6 @@ This template uses **pre-commit hooks** to automatically check your code before
| [Biome](https://biomejs.dev/) | Formats JSON/JSONC files |
| [pyproject-fmt](https://github.com/tox-dev/pyproject-fmt) | Formats `pyproject.toml` |

**Notebook outputs** are handled separately by an [nbstripout](https://github.com/kynan/nbstripout)
git *filter* (not a pre-commit hook), set up by `pixi run install-hooks`. The filter
strips outputs from the committed copy while leaving them in your working tree, so
your notebooks stay rendered locally but git history stays clean. CI fails the build
if a notebook with outputs ever lands in the repo (a clone that skipped
`install-hooks`), via `nbstripout --verify`.

Hooks run automatically on `git commit`. To run manually:

```bash
Expand Down
90 changes: 90 additions & 0 deletions REVIEW_GUIDE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
# Review Guide

PR review playbook for **review agents running on GitHub**. Imperative voice.

**Scope: review only.** Comment and suggest. Do not push commits or apply fixes.

> Everything above "This repo" comes from the analysis template, which takes it from the analysis-workflow
> plugin; change it there and pull it with `cruft update`. Reviewers cannot load the plugin, which is why this
> copy exists.

The code has usually already run. That is not a reason to lower the bar: requiring a rerun in better shape is a
normal outcome. Say what to redo, not just what was wrong.

## Data governance

Every write to a shared object (a working object under `data/*/processed/`, a SpatialData store, anything other
code loads):

- **Accumulate by addition.** Adding keys to freshly re-read state is commutative, so concurrent sessions can't
lose each other's work. Removal isn't: that means a new dated copy.
- **Writing back to a shared object needs sign-off on that specific diff.**

## Priorities

### 1. Writes that lose data, or silently do nothing

A write to a working object goes through `commit_adata`, gated behind a dry run. Flag:

- an in-memory AnnData written back over a working object (`write_zarr`/`write_h5ad` onto its path): it erases
whatever other sessions added since it was read;
- a recomputed key committed under an existing name: `commit_adata` writes only keys not yet on disk, so it is a
no-op and the PR reports a result the object doesn't hold. Use a new key or a new dated copy;
- a replaced SpatialData image, labels, points or shapes element, where a new element name would do.

### 2. Promotions, which leave no diff

`data/` is gitignored, so an object written into `data/<dataset>/processed/` appears nowhere in the diff; the task
README's `Write-back` section is the only trace. The right subdirectory is not guessable. Check that the PR *says*
where it went and that it was agreed, while moving it is still cheap.

### 3. The task contract

Agent work belongs in a task directory `analysis/<topic>/.../<name>_vN/` with a `README.md` naming real inputs,
outputs and write-back keys. Outputs go through `task_paths(__file__)` to the task's own `results/`, `reports/`,
`figures/`, `outputs/`. Writes to `data/<dataset>/results/` or the central `figures/` are the human lane and are a
finding.

Every key committed to a shared object carries the task's version suffix **and** appears in that README: an
unrecorded key can't be traced back, an unsuffixed one collides with the next version.

### 4. Scientific correctness

The point of the repo. Check that the claim in the PR body follows from what the code computes: the right cells,
the right grouping, a control where one is needed, and a comparison not confounded by condition, time point,
sample or batch. A number that reaches a figure is worth more scrutiny than anything below.

### 5. Reinvention

New code needs a reason to exist. Look outward before accepting it: scverse and the Python ecosystem, then the
repo's sibling code packages, then earlier task directories; a near-match found by grep counts. Name the
candidate and what's wrong with it rather than concluding nothing fits.

The acute case is a helper defined twice in one task or copied between tasks: copies drift, and when they compute
a quantity compared *across* tasks, the result is a wrong conclusion. Copying figure code between versions of one
task is fine.

### 6. Conciseness

Prose is reviewed like code. READMEs, docstrings, comments and the PR body: no restating the diff, no filler
preamble, no exhaustive caveats. Padding rots the same way dead code does, and these docs are load-bearing for
the next agent.

## Do not report

- Anything the CI linters already gate: formatting, imports, line length.
- Lock files, and anything under `data/`.
- Committed notebook outputs: review changed cells' code, never the output blobs.
- Absolute paths built from `main_checkout()`: task outputs are anchored there deliberately.
- Worktree, temp-dir and cwd-relative `data/`/`figures/` path literals in `analysis/` code: a hook blocks them
before the file is written.

## Drift older than the diff

Parts of a repo may predate these rules. When a PR touches such a file, report the drift once at the lowest
severity and don't block; it gets fixed when a task next touches it. Search the open issues before filing one.

## This repo

<!-- Repo-specific rules go here: its shared stores and their write helpers, its companion packages, its
confounders. -->
6 changes: 1 addition & 5 deletions analysis/ML-2026-01-27_demo_scRNA_workflow.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -497,11 +497,7 @@
"id": "37",
"metadata": {},
"outputs": [],
"source": [
"output_path = FilePaths.EXAMPLE_DATASET / \"processed\" / \"pbmc3k_processed.h5ad\"\n",
"adata.write(output_path)\n",
"print(f\"Saved to: {output_path}\")"
]
"source": "output_path = FilePaths.EXAMPLE_DATASET / \"processed\" / \"pbmc3k_processed.zarr\"\nadata.write_zarr(output_path)\nprint(f\"Saved to: {output_path}\")"
},
{
"cell_type": "markdown",
Expand Down
4 changes: 1 addition & 3 deletions analysis/XX-2026-01-27_sample_notebook.ipynb
Original file line number Diff line number Diff line change
Expand Up @@ -127,9 +127,7 @@
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Store data in `data/<dataset>/` with subfolders `raw/`, `processed/`, `resources/`, `results/`. Access paths via `FilePaths`:"
]
"source": "Store data in `data/<dataset>/` (layout in the README) and access paths via `FilePaths`:"
},
{
"cell_type": "code",
Expand Down
6 changes: 2 additions & 4 deletions data/example_dataset/README.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,4 @@
# Dataset structure

- `raw`: Raw state of the data we received.
- `processed`: Processed / intermediate data.
- `resources`: Reference data, gene sets, annotations.
- `results`: Any results we compute for this dataset.
`raw/`, `resources/`, `processed/` and `results/`, as described under *Data and notebook
conventions* in the top-level `README.md`.
Loading
Loading