Skip to content

Repository files navigation

CuratorKIT

Data curation and synthetic data generation for LLM post-training.
Ingest from any source, generate with any LLM, verify every sample against its origin,
recover what fails, and export trainer-ready datasets with full provenance.

PyPI version Python 3.11+ License: LSAL v1.1 (source-available) Documentation

Documentation · Quickstart · Tutorials · Ecosystem · Contributing


Why CuratorKIT

Synthetic data is the fastest way to build a post-training dataset. It is also the fastest way to poison a model: generated samples hallucinate beyond their source material, drift off-distribution, duplicate each other, and quietly carry secrets, PII, and toxicity into training runs. Most teams find out after fine-tuning.

CuratorKIT treats dataset construction as a pipeline with quality gates, not a script with a prompt. Every generated answer is checked for grounding against the exact source chunk it came from. Rejected samples get diagnosed rather than discarded, and the fixable ones are repaired. Every run emits a provenance manifest, a dataset card, and checksums, so any sample can be audited back to its source.

Key features

  • Grounded hallucination gate. Verifies each generated answer against the source passage it was generated from, not against the judge model's general knowledge.
  • Reward and diversity gates. Multi-dimension LLM-judge scoring plus embedding-based filtering against low-quality and near-duplicate samples.
  • Adaptive recovery. An inline diagnostic probe classifies each rejection into a failure-mode taxonomy and repairs the recoverable ones.
  • Data hygiene. Secrets detection, PII pseudonymization (Presidio), and toxicity filtering as pipeline stages.
  • Any source. JSONL, JSON, CSV, Parquet, HuggingFace datasets, and PDFs with layout-aware parsing. Multi-source runs with per-source field mapping.
  • Eight generation tasks. QA, preference pairs, GRPO rollouts, multi-turn, Evol-Instruct, chain-of-thought, and adversarial variants, on any LiteLLM-compatible API or local Ollama/vLLM.
  • Trainer-ready exports. Alpaca, ShareGPT, chat messages, DPO, GRPO, and PPO formats with train/val/test splits, consumed directly by TRL and AlignTune.
  • Provenance by default. Every run writes manifest.json, rejected.jsonl with structured reasons, dataset_card.md, lexsi_provenance.json, and SHA-256 checksums.txt.

How it works

flowchart LR
    A["Ingest<br/>PDF · JSONL · CSV<br/>Parquet · HF Hub"] --> B["Clean + dedup<br/>exact · minhash · embedding"]
    B --> C["Hygiene<br/>secrets · PII · toxicity"]
    C --> D["Generate<br/>QA · DPO · GRPO · CoT"]
    D --> E{"Quality gates<br/>hallucination · reward · diversity"}
    E -->|pass| F["Export<br/>Alpaca · ShareGPT · Messages · DPO · GRPO · PPO"]
    E -->|reject| G["Adaptive recovery<br/>diagnose · repair · re-gate"]
    G -->|recovered| E
    G -->|unrecoverable| H["rejected.jsonl<br/>+ diagnostics"]
    F --> I["manifest · dataset card · checksums"]
Loading

Install

pip install "curatorkit[all]"        # connectors + generation + embedding + hygiene (recommended)
pip install curatorkit               # core only: cleaning and dedup

Requires Python 3.11+ on Linux, macOS, or Windows.

More install options
pip install "curatorkit[generation]"       # LLM generation only
pip install "curatorkit[generation-full]"  # generation + embedding + FAISS
pip install "curatorkit[connectors]"       # pyarrow + datasets (no LLM)
pip install "curatorkit[hygiene]"          # secrets / PII / toxicity gates
pip install "curatorkit[pdf]"              # layout-aware PDF parsing (MinerU)

# From source
pip install "curatorkit[all] @ git+https://github.com/Lexsi-Labs/CuratorKIT.git"

The installation guide has the full extras table.

Quickstart

Clean and deduplicate an existing dataset (no LLM, no API key):

from curatorkit import Curator, CuratorConfig

result = Curator(CuratorConfig(
    dataset        = {"name": "tatsu-lab/alpaca", "max_samples": 2000},
    dedup          = "minhash",
    clean          = True,
    export_formats = ["alpaca", "sharegpt"],
    output_dir     = "output/clean",
)).run()

result.print_summary()

HF Hub sources need the connectors extra (included in all); local JSONL/CSV files run on the core install.

Generate gated synthetic QA data from a document:

export OPENAI_API_KEY=sk-...   # any LiteLLM backend works; local Ollama/vLLM too
result = Curator(CuratorConfig(
    dataset                 = "handbook.pdf",          # needs the [pdf] extra
    llm_model               = "openai/gpt-4o-mini",
    generation_task         = "qa",
    num_questions           = 3,
    hallucination_threshold = 0.7,                     # grounding gate
    reward_threshold        = 0.7,                     # LLM-judge gate
    export_formats          = ["alpaca", "sharegpt"],
    output_dir              = "output/qa",
)).run()

Or run a declarative YAML pipeline from the CLI:

curatorkit run examples/quickstart/pipeline.yaml --output-dir output/

Every run writes manifest.json, rejected.jsonl, dataset_card.md, lexsi_provenance.json, and checksums.txt alongside the export files. The quickstart example runs end to end without an API key.

The output folder is also a Hugging Face dataset: its README.md maps each export to a config (sft_alpaca, sft_sharegpt, sft_messages, dpo, grpo, ppo, corpus) with one split per output_split directory, or train:

from datasets import load_dataset
chat = load_dataset("output/qa", "sft_messages")  # DatasetDict: train (+ validation, ... with output_split)

Cohere Aya as generator or judge

Any LiteLLM model string works, so Aya can generate the data, judge it, or both. Three ways to reach it:

from curatorkit import Curator, CuratorConfig

result = Curator(CuratorConfig(
    dataset                 = "handbook.pdf",
    # 1. Cohere API (export COHERE_API_KEY=...)
    llm_model               = "cohere_chat/c4ai-aya-expanse-32b",
    # 2. Ollama:  llm_model = "ollama/aya-expanse"
    # 3. A local vLLM / `transformers serve` endpoint (OpenAI-compatible):
    #    llm_model = "openai/CohereLabs/aya-expanse-8b", llm_api_base = "http://localhost:8000/v1"
    judge_llm_model         = "cohere_chat/c4ai-aya-expanse-8b",   # optional; defaults to llm_model
    generation_task         = "qa",
    hallucination_threshold = 0.7,
    reward_threshold        = 0.7,
    min_tokens              = 3,          # short multilingual rows; CJK/Thai count per character
    output_dir              = "output/",
)).run()

result.push_to_hub("your-org/aya-qa", export_file="sft_alpaca.jsonl")   # private by default; needs [hf]
  • Tiny Aya (e.g. CohereLabs/tiny-aya-global) is gated on the Hugging Face Hub: request access on the model page and log in with hf auth login before serving it with route 3. It is not in the Ollama library.
  • Aya Vision served through route 3 needs a chat template that accepts list-type content; CuratorKIT sends text only.
  • Small judges often fail to return JSON. A failed judge call or an unparseable score rejects the sample (judge_error:<type> in rejected.jsonl). Set judge_on_error="pass" to keep such samples instead.

What it generates

generation_task Output task type Use for
qa instruction_following SFT on question-answering
preference preference DPO training
grpo grpo GRPO training
multiturn conversational Multi-turn SFT
evol instruction_following Harder instruction variants
cot instruction_following Chain-of-thought reasoning
adversarial_preference preference Robustness DPO
adversarial_qa instruction_following Hallucination stress-test data

Output files

File Always written Contents
manifest.json ✓ Config hash, per-stage counts, rejection breakdown
rejected.jsonl ✓ All rejected samples with structured reason strings
dataset_card.md ✓ Human-readable run summary
README.md ✓* The same card with configs: YAML, so load_dataset(output_dir, "<config>") works
lexsi_provenance.json ✓ Library, version, inputs, and generator/judge model ids (also manifest["provenance"])
checksums.txt ✓ SHA-256 for all output files
sft_alpaca.jsonl optional Alpaca-format SFT data (instruction, input, output)
sft_sharegpt.jsonl optional ShareGPT conversations (from/value)
sft_messages.jsonl optional Chat messages (role/content), as TRL's SFTTrainer expects
dpo.jsonl optional DPO preference pairs
grpo.jsonl optional GRPO group rollouts
ppo.jsonl optional PPO prompt-only format
diagnostic_summary.json optional Failure mode counts, recovery rate (when probe active)

* An existing README.md that CuratorKIT did not write is left untouched (with a warning), and write_hf_readme=False skips it. Use a dedicated output_dir to get a loadable folder.

Handing off to AlignTune / SafeTune / AuditKIT

Pass the output folder and a config name; no conversion step. Use output_split={"train": 0.9, "validation": 0.1} to get a validation split.

# AlignTune (Track 1): SFT on the chat messages, or DPO on the preference pairs
from aligntune.core.backend_factory import create_sft_trainer
create_sft_trainer(model_name="CohereLabs/aya-expanse-8b", dataset_name="output/", config_name="sft_messages", backend="trl").train()
# SafeTune (Track 2): the curated set is the fine-tuning data for a harden method
from datasets import load_dataset
from safetune.runner import harden
harden.SafeGradTrainer(model, tokenizer).train(load_dataset("output/", "sft_messages", split="train"), safety_dataset=safety_dataset)
# AuditKIT: evaluate on the held-out split
from auditkit.loaders import load_hf
samples = load_hf("output/", name="sft_alpaca", split="validation", input_col="instruction", target_col="output")

Each library records lexsi_provenance.json from this folder in its own outputs, so the lineage follows the model through the track.

Documentation

Full documentation lives at lexsi-labs.github.io/CuratorKIT.

Section Contents
Getting started Installation, quickstart, reading the output
Guides Data sources, generation, quality gates, recovery, hygiene, exporters, customisation
Configuration reference Every CuratorConfig parameter
CLI & YAML curatorkit run, all flags, the YAML pipeline schema
API reference Generated from the source docstrings
Architecture Base classes, contracts, provenance model

Tutorials

Nine runnable notebooks cover the feature set end-to-end.

# Notebook Links
01 Generate an SFT dataset from a PDF Open In Colab
02 Generate DPO preference pairs Open In Colab
03 Generate GRPO rollouts Open In Colab
04 Ingest multiple sources Open In Colab
05 Clean and deduplicate a dataset Open In Colab
06 Adaptive recovery Open In Colab
07 Adversarial generation Open In Colab
08 Data hygiene pipeline Open In Colab
09 Filtered vs unfiltered fine-tuning Open In Colab

Full descriptions and prerequisites are in the tutorials index.

Contributing

The connector, generator, gate, and exporter layers are designed as plugin points. Read the contributing guide and the architecture reference, then open an issue or PR. Questions go to GitHub Issues.

Support

Citation

If you use CuratorKIT in your research, please cite it (see CITATION.cff):

@software{curatorkit2026,
  author    = {Bhattacharjee, Soham and Sharma, Karun and Sankarapu, Vinay Kumar and Seth, Pratinav},
  title     = {CuratorKIT: Data Curation and Synthetic Data Generation for LLM Post-Training},
  year      = {2026},
  publisher = {Lexsi Labs},
  url       = {https://github.com/Lexsi-Labs/CuratorKIT}
}

@misc{bhattacharjee2026curatorkitdatacuration,
      title={CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training}, 
      author={Soham Bhattacharjee and Karun Sharma and Vinay Kumar Sankarapu and Pratinav Seth},
      year={2026},
      eprint={2606.21631},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.21631}, 
}

@misc{bhattacharjee2026provenancegroundedgatingadaptiverecovery,
      title={Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation}, 
      author={Soham Bhattacharjee and Karun Sharma and Vinay Kumar Sankarapu and Pratinav Seth},
      year={2026},
      eprint={2606.11127},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.11127}, 
}

License

LSAL v1.1; see LICENSE. Free for research, education, and non-commercial use. Commercial use requires a separate license — contact support@lexsi.ai. The optional pdf extra installs MinerU, which is licensed AGPL-3.0. Install it only if that suits your use.


Contact

Lexsi Labs

https://www.lexsi.ai

Paris 🇫🇷 · Mumbai 🇮🇳 · London 🇬🇧

Ecosystem

CuratorKIT is part of the Lexsi Labs open-source stack:

  • AlignTune fine-tunes with the data you curate here. CuratorKIT's Alpaca, DPO, GRPO, and PPO exports are AlignTune's native input formats.
  • TabTune is a unified library for tabular foundation models.
  • DLBacktrace and xai_evals cover model interpretability and explanation evaluation.

About

CuratorKIT : Data curation and synthetic data generation for LLM post-training.

Resources

Code of conduct

Contributing

Security policy

Stars

29 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages