Data curation and synthetic data generation for LLM post-training.
Ingest from any source, generate with any LLM, verify every sample against its origin,
recover what fails, and export trainer-ready datasets with full provenance.
Documentation · Quickstart · Tutorials · Ecosystem · Contributing
Synthetic data is the fastest way to build a post-training dataset. It is also the fastest way to poison a model: generated samples hallucinate beyond their source material, drift off-distribution, duplicate each other, and quietly carry secrets, PII, and toxicity into training runs. Most teams find out after fine-tuning.
CuratorKIT treats dataset construction as a pipeline with quality gates, not a script with a prompt. Every generated answer is checked for grounding against the exact source chunk it came from. Rejected samples get diagnosed rather than discarded, and the fixable ones are repaired. Every run emits a provenance manifest, a dataset card, and checksums, so any sample can be audited back to its source.
- Grounded hallucination gate. Verifies each generated answer against the source passage it was generated from, not against the judge model's general knowledge.
- Reward and diversity gates. Multi-dimension LLM-judge scoring plus embedding-based filtering against low-quality and near-duplicate samples.
- Adaptive recovery. An inline diagnostic probe classifies each rejection into a failure-mode taxonomy and repairs the recoverable ones.
- Data hygiene. Secrets detection, PII pseudonymization (Presidio), and toxicity filtering as pipeline stages.
- Any source. JSONL, JSON, CSV, Parquet, HuggingFace datasets, and PDFs with layout-aware parsing. Multi-source runs with per-source field mapping.
- Eight generation tasks. QA, preference pairs, GRPO rollouts, multi-turn, Evol-Instruct, chain-of-thought, and adversarial variants, on any LiteLLM-compatible API or local Ollama/vLLM.
- Trainer-ready exports. Alpaca, ShareGPT, chat messages, DPO, GRPO, and PPO formats with train/val/test splits, consumed directly by TRL and AlignTune.
- Provenance by default. Every run writes
manifest.json,rejected.jsonlwith structured reasons,dataset_card.md,lexsi_provenance.json, and SHA-256checksums.txt.
flowchart LR
A["Ingest<br/>PDF · JSONL · CSV<br/>Parquet · HF Hub"] --> B["Clean + dedup<br/>exact · minhash · embedding"]
B --> C["Hygiene<br/>secrets · PII · toxicity"]
C --> D["Generate<br/>QA · DPO · GRPO · CoT"]
D --> E{"Quality gates<br/>hallucination · reward · diversity"}
E -->|pass| F["Export<br/>Alpaca · ShareGPT · Messages · DPO · GRPO · PPO"]
E -->|reject| G["Adaptive recovery<br/>diagnose · repair · re-gate"]
G -->|recovered| E
G -->|unrecoverable| H["rejected.jsonl<br/>+ diagnostics"]
F --> I["manifest · dataset card · checksums"]
pip install "curatorkit[all]" # connectors + generation + embedding + hygiene (recommended)
pip install curatorkit # core only: cleaning and dedupRequires Python 3.11+ on Linux, macOS, or Windows.
More install options
pip install "curatorkit[generation]" # LLM generation only
pip install "curatorkit[generation-full]" # generation + embedding + FAISS
pip install "curatorkit[connectors]" # pyarrow + datasets (no LLM)
pip install "curatorkit[hygiene]" # secrets / PII / toxicity gates
pip install "curatorkit[pdf]" # layout-aware PDF parsing (MinerU)
# From source
pip install "curatorkit[all] @ git+https://github.com/Lexsi-Labs/CuratorKIT.git"The installation guide has the full extras table.
Clean and deduplicate an existing dataset (no LLM, no API key):
from curatorkit import Curator, CuratorConfig
result = Curator(CuratorConfig(
dataset = {"name": "tatsu-lab/alpaca", "max_samples": 2000},
dedup = "minhash",
clean = True,
export_formats = ["alpaca", "sharegpt"],
output_dir = "output/clean",
)).run()
result.print_summary()HF Hub sources need the connectors extra (included in all); local JSONL/CSV files run on the core install.
Generate gated synthetic QA data from a document:
export OPENAI_API_KEY=sk-... # any LiteLLM backend works; local Ollama/vLLM tooresult = Curator(CuratorConfig(
dataset = "handbook.pdf", # needs the [pdf] extra
llm_model = "openai/gpt-4o-mini",
generation_task = "qa",
num_questions = 3,
hallucination_threshold = 0.7, # grounding gate
reward_threshold = 0.7, # LLM-judge gate
export_formats = ["alpaca", "sharegpt"],
output_dir = "output/qa",
)).run()Or run a declarative YAML pipeline from the CLI:
curatorkit run examples/quickstart/pipeline.yaml --output-dir output/Every run writes manifest.json, rejected.jsonl, dataset_card.md, lexsi_provenance.json, and checksums.txt alongside the export files. The quickstart example runs end to end without an API key.
The output folder is also a Hugging Face dataset: its README.md maps each export to a config (sft_alpaca, sft_sharegpt, sft_messages, dpo, grpo, ppo, corpus) with one split per output_split directory, or train:
from datasets import load_dataset
chat = load_dataset("output/qa", "sft_messages") # DatasetDict: train (+ validation, ... with output_split)Any LiteLLM model string works, so Aya can generate the data, judge it, or both. Three ways to reach it:
from curatorkit import Curator, CuratorConfig
result = Curator(CuratorConfig(
dataset = "handbook.pdf",
# 1. Cohere API (export COHERE_API_KEY=...)
llm_model = "cohere_chat/c4ai-aya-expanse-32b",
# 2. Ollama: llm_model = "ollama/aya-expanse"
# 3. A local vLLM / `transformers serve` endpoint (OpenAI-compatible):
# llm_model = "openai/CohereLabs/aya-expanse-8b", llm_api_base = "http://localhost:8000/v1"
judge_llm_model = "cohere_chat/c4ai-aya-expanse-8b", # optional; defaults to llm_model
generation_task = "qa",
hallucination_threshold = 0.7,
reward_threshold = 0.7,
min_tokens = 3, # short multilingual rows; CJK/Thai count per character
output_dir = "output/",
)).run()
result.push_to_hub("your-org/aya-qa", export_file="sft_alpaca.jsonl") # private by default; needs [hf]- Tiny Aya (e.g.
CohereLabs/tiny-aya-global) is gated on the Hugging Face Hub: request access on the model page and log in withhf auth loginbefore serving it with route 3. It is not in the Ollama library. - Aya Vision served through route 3 needs a chat template that accepts list-type
content; CuratorKIT sends text only. - Small judges often fail to return JSON. A failed judge call or an unparseable score rejects the sample (
judge_error:<type>inrejected.jsonl). Setjudge_on_error="pass"to keep such samples instead.
generation_task |
Output task type | Use for |
|---|---|---|
qa |
instruction_following |
SFT on question-answering |
preference |
preference |
DPO training |
grpo |
grpo |
GRPO training |
multiturn |
conversational |
Multi-turn SFT |
evol |
instruction_following |
Harder instruction variants |
cot |
instruction_following |
Chain-of-thought reasoning |
adversarial_preference |
preference |
Robustness DPO |
adversarial_qa |
instruction_following |
Hallucination stress-test data |
| File | Always written | Contents |
|---|---|---|
manifest.json |
✓ | Config hash, per-stage counts, rejection breakdown |
rejected.jsonl |
✓ | All rejected samples with structured reason strings |
dataset_card.md |
✓ | Human-readable run summary |
README.md |
✓* | The same card with configs: YAML, so load_dataset(output_dir, "<config>") works |
lexsi_provenance.json |
✓ | Library, version, inputs, and generator/judge model ids (also manifest["provenance"]) |
checksums.txt |
✓ | SHA-256 for all output files |
sft_alpaca.jsonl |
optional | Alpaca-format SFT data (instruction, input, output) |
sft_sharegpt.jsonl |
optional | ShareGPT conversations (from/value) |
sft_messages.jsonl |
optional | Chat messages (role/content), as TRL's SFTTrainer expects |
dpo.jsonl |
optional | DPO preference pairs |
grpo.jsonl |
optional | GRPO group rollouts |
ppo.jsonl |
optional | PPO prompt-only format |
diagnostic_summary.json |
optional | Failure mode counts, recovery rate (when probe active) |
* An existing README.md that CuratorKIT did not write is left untouched (with a warning), and write_hf_readme=False skips it. Use a dedicated output_dir to get a loadable folder.
Pass the output folder and a config name; no conversion step. Use output_split={"train": 0.9, "validation": 0.1} to get a validation split.
# AlignTune (Track 1): SFT on the chat messages, or DPO on the preference pairs
from aligntune.core.backend_factory import create_sft_trainer
create_sft_trainer(model_name="CohereLabs/aya-expanse-8b", dataset_name="output/", config_name="sft_messages", backend="trl").train()# SafeTune (Track 2): the curated set is the fine-tuning data for a harden method
from datasets import load_dataset
from safetune.runner import harden
harden.SafeGradTrainer(model, tokenizer).train(load_dataset("output/", "sft_messages", split="train"), safety_dataset=safety_dataset)# AuditKIT: evaluate on the held-out split
from auditkit.loaders import load_hf
samples = load_hf("output/", name="sft_alpaca", split="validation", input_col="instruction", target_col="output")Each library records lexsi_provenance.json from this folder in its own outputs, so the lineage follows the model through the track.
Full documentation lives at lexsi-labs.github.io/CuratorKIT.
| Section | Contents |
|---|---|
| Getting started | Installation, quickstart, reading the output |
| Guides | Data sources, generation, quality gates, recovery, hygiene, exporters, customisation |
| Configuration reference | Every CuratorConfig parameter |
| CLI & YAML | curatorkit run, all flags, the YAML pipeline schema |
| API reference | Generated from the source docstrings |
| Architecture | Base classes, contracts, provenance model |
Nine runnable notebooks cover the feature set end-to-end.
Full descriptions and prerequisites are in the tutorials index.
The connector, generator, gate, and exporter layers are designed as plugin points. Read the contributing guide and the architecture reference, then open an issue or PR. Questions go to GitHub Issues.
- Documentation: lexsi-labs.github.io/CuratorKIT
- GitHub Issues: github.com/Lexsi-Labs/CuratorKIT/issues
- Email: pratinav.seth@lexsi.ai
If you use CuratorKIT in your research, please cite it (see CITATION.cff):
@software{curatorkit2026,
author = {Bhattacharjee, Soham and Sharma, Karun and Sankarapu, Vinay Kumar and Seth, Pratinav},
title = {CuratorKIT: Data Curation and Synthetic Data Generation for LLM Post-Training},
year = {2026},
publisher = {Lexsi Labs},
url = {https://github.com/Lexsi-Labs/CuratorKIT}
}
@misc{bhattacharjee2026curatorkitdatacuration,
title={CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training},
author={Soham Bhattacharjee and Karun Sharma and Vinay Kumar Sankarapu and Pratinav Seth},
year={2026},
eprint={2606.21631},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2606.21631},
}
@misc{bhattacharjee2026provenancegroundedgatingadaptiverecovery,
title={Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation},
author={Soham Bhattacharjee and Karun Sharma and Vinay Kumar Sankarapu and Pratinav Seth},
year={2026},
eprint={2606.11127},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2606.11127},
}LSAL v1.1; see LICENSE. Free for research, education, and non-commercial use. Commercial use requires a separate license — contact support@lexsi.ai. The optional pdf extra installs MinerU, which is licensed AGPL-3.0. Install it only if that suits your use.
CuratorKIT is part of the Lexsi Labs open-source stack:
- AlignTune fine-tunes with the data you curate here. CuratorKIT's Alpaca, DPO, GRPO, and PPO exports are AlignTune's native input formats.
- TabTune is a unified library for tabular foundation models.
- DLBacktrace and xai_evals cover model interpretability and explanation evaluation.