Research files prepared through a readable YAML protocol, with checked tables and execution receipts.
Wrangle brings inspection, cleaning, matching, reshaping, and validation into one preparation protocol. Declare the research decisions, prepare the table, and keep the evidence needed to review the result and reuse the protocol. Researchers, notebooks, and agents use the same native Polars engine.
Wrangle owns data preparation, protocol checks, and retained evidence. Researchers own observation identity, units, missingness, exclusions, metadata matches, and statistical choices. Wrangle does not infer those decisions or certify their scientific suitability. It does not build or fit models. All transformations and data checks compile to Polars expressions and LazyFrame plans; Python handles orchestration, files, and receipts.
| Research task | Supported capability |
|---|---|
| Understand incoming files | Inspect columns, missingness, example rows, summaries, and schema changes |
| Prepare instrument measurements | Preserve identifiers; explicitly parse numbers, dates, missing codes, and text |
| Match samples and metadata | Joins with enforced cardinality, ordering, and unmatched-observation policies |
| Combine batches or reshape observations | Concatenation, pivoting, unpivoting, nested fields, and declared output identity |
| Apply a research protocol | Category mappings, derived fields, unit conversions, and recorded exclusions |
| Prepare longitudinal or grouped data | Group summaries, temporal matching, lag, and rolling calculations |
| Reuse reference-batch decisions | Frozen imputation and standardization parameters; declared sampling and partitions |
| Check and retain an analysis table | Output contracts, readable reports, reusable YAML, and hash receipts |
Use the recipe manual to choose a workflow and the operation catalog for exact arguments, preconditions, and stable failures.
Use Python 3.10 or later. Install Wrangle from PyPI:
python -m pip install --upgrade wrangleInstallation supplies Polars and the strict YAML parser. Version 1.0 replaces the
historical 0.x interface; follow the migration guide
before upgrading existing code.
For Excel, install python -m pip install 'wrangle[excel]' and select
a worksheet explicitly. Choose new directories for examples and outputs;
existing destinations are rejected.
Run this complete example in a terminal:
wrangle example my-study
cd my-study
wrangle inspect samples.csv
wrangle prepare recipe.yaml --source measurements=samples.csv --source metadata=metadata.csv --output prepared
wrangle inspect preparedExpected result: three input samples become two accepted samples, with metadata attached and mass converted from milligrams to grams:
| sample_id | mass_g | qc | group |
|---|---|---|---|
| 001 | 1.0 | pass | control |
| 003 | 3.0 | pass | treated |
Sample 002 is excluded by the example's declared instrument quality rule.
Open prepared/report.txt to review the checks and exclusion. The commented
recipe.yaml explains the supplied decisions; its quality rule, units, and limits
belong to this example and must be adapted to your study.
From the my-study directory:
import wrangle as wr
profile = wr.inspect("samples.csv")
print(profile.summary())
result = wr.prepare(
{"measurements": "samples.csv", "metadata": "metadata.csv"},
"recipe.yaml",
output="prepared-python",
)
print(result.summary())
result.data # A Polars table with the two accepted samples.inspect observes without changing the source. CSV fields start as text,
preserving identifiers such as 001; the recipe explicitly converts measurements.
prepare returns checked data and a receipt. A one-source recipe can use
wr.prepare("samples.csv", "recipe.yaml"); Python callers may also supply a mapping.
wrangle start measurements.csv --metadata metadata.csv --output my-protocolOmit --metadata metadata.csv for one file. The guide asks what a row represents,
which fields identify it, how numbers and units should be represented, and what
missing values, exclusions, and metadata matches mean. Enter column numbers or
names; press Enter if you are unsure. Your answers become a commented YAML protocol.
Unanswered decisions block preparation. A complete draft is ready to run checks;
preparation validates the data. Run the command printed by the guide and review
the resulting report. Reuse the same recipe for the next batch.
For Excel, supply --sheet "Sheet name" and, when needed,
--metadata-sheet "Sheet name". See Getting started.
| File | Retained evidence |
|---|---|
data.parquet |
Checked analysis table |
recipe.yaml |
Resolved, reusable preparation protocol |
report.txt |
Readable summary of checks, changes, exclusions, and unresolved declarations |
receipt.json |
Input and output hashes, protocol identity, checks, exclusions, and fitted parameters |
Python's output="new-directory" and the CLI's --output new-directory save
these files together. Use result.write("new-directory") to save an unsaved
Python result. Sources remain unchanged; existing outputs are never overwritten.
A saved preparation directory can be reused as a source after evidence verification.
The same protocol serves shell scripts and agents:
wrangle inspect samples.csv --json
wrangle catalog join --json
wrangle prepare recipe.yaml --source measurements=samples.csv --source metadata=metadata.csv --jsonThe CLI is readable by default; --json supplies structured results and errors.
Agents use wrangle start ... --json for the same questions without prompts and
--answers answers.yaml for researcher-approved decisions. Preparation uses
identical checks and evidence in Python and both CLI presentations.
Recipe files are YAML only (.yaml or .yml); internal receipts and catalogs use
JSON. Comments and formatting do not change protocol identity.
Passing checks establishes conformity to your declared protocol. It does not
establish the scientific correctness of a chosen rule. Hash receipts check
consistency with retained evidence; they do not establish authorship or constitute
an independent audit. Unresolved guided decisions stop execution with
UNRESOLVED_PROTOCOL; answer them rather than removing the blockers.
Preparation accepts local CSV, TSV, Parquet, Arrow IPC, NDJSON, explicitly selected Excel sheets, native Polars tables, and verified preparation bundles. It never uploads data or accesses the network. Recipes cannot execute arbitrary callbacks. Read the security boundaries and recipe contracts and recovery.
For large local files, use --execution disk --output new-directory, or
wr.prepare(sources, "recipe.yaml", execution="disk", output="new-directory").
Complete input snapshots, intermediates, and evidence stay on disk; result.data
is a Polars LazyFrame. Global joins, sorting, and grouped calculations can still
need substantial RAM. Inspection and guided start remain eager.
See disk execution.
| Job | Start here |
|---|---|
| Prepare your own files with guided decisions | Getting started |
| Try the complete human and agent example | Verified first study |
| Keep and reuse a research batch | Research batch |
| Join, reshape, or summarize observations | Preparation workflows |
| Reuse fitted reference parameters | Frozen parameters |
| Prepare large files on disk | Disk workflow |
| Select an operation and its preconditions | Operation catalog |
| Integrate Python or an agent | Agent entrypoint and recipe manual |
| Run commands and recover from errors | Command line |
| Replace a historical method | Migration |
Installed agents start at Path(wrangle.__file__).parent / "AGENTS.md";
the installed manual and generated operation contracts live beside it under docs/.
The catalog is generated from the engine, and shipped examples are executed in
release checks. Stable WrangleError codes and details guide recovery without
substituting a scientific method. Dedicated Wrangle agent tools are unnecessary.
Start with CONTRIBUTING.md for setup, checks, and review. Propose work through Wrangle issues. Governance, conduct, the roadmap, and release notes describe project maintenance; architecture describes the engine.
Use Wrangle issues for bug reports, feature requests, and usage questions. Include a minimal YAML protocol, Wrangle and Python versions, operating system, and the error code and details. Use a small shareable fixture when a report needs source data.
Report suspected vulnerabilities privately through GitHub Security Advisories. Use SECURITY.md for supported versions and response policy, and release verification to check artifact identity. OpenSSF evidence records the published Silver answers, distinguishing measured checks from owner attestations.
Identify Wrangle, its version, and the retained preparation protocol and receipt when citing a research workflow. Link the corresponding release to identify the software used. No DOI is supplied.