Skip to content

feat: Cinto Kernel — pipeline runtime for local AI - #1

Merged
Jshebb merged 43 commits into
masterfrom
feat/cinto-kernel-architecture
May 15, 2026
Merged

Jshebb merged 43 commits into
masterfrom
feat/cinto-kernel-architecture

Conversation

@Jshebb

@Jshebb Jshebb commented May 14, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds Cinto Kernel — a pipeline runtime that turns small local models into reliable coding agents by giving each stage one bounded job instead of asking the model to do everything at once.

What's in this PR

  • Kernel runtime — 5-stage pipeline: Interpret → Locate → Hypothesize → Patch → Report. Each stage gets a slot-specific context pack, one CRP template, and retries independently.
  • Syscall surface — index_repo, search, read_range, read_around, list_symbols, context_pack — all with hard output caps so models aren't overwhelmed
  • Patch applier — parses <EDIT> CRP blocks and @@ terse format; replace_function falls back to append for new functions; indentation-based scoping for Python
  • Local memory — .cinto/memory/decisions.md written after each HYPOTHESIZE stage; loaded into future runs for context across sessions
  • TUI integration — F5 / /kernel toggles between Core (conversational) and Kernel (pipeline) modes
  • Small model support — lm-studio-small + qwen3 presets, no_think config field, think-block stripping, plain-text fallback, per-stage one-shot CRP examples
  • Eval infrastructure — cinto batch --kernel, context window auto-detection, cinto use-preset, cinto presets
  • Setup fix — config now saves on Esc/Tab exit, not only on explicit "Save & Start"
  • 114 passing tests, pipeline architecture documented in wiki

Eval results (baseline tasks, no fine-tuning)

Python (Qwen3.5-9B) — 4/5 tasks passing (80%)

Task Result Change applied
py_bugfix_01 PASS find_max off-by-one fixed
py_bugfix_02 PASS calculate_average empty list handled
py_feature_01 PASS getters added to User class (full replace)
py_refactor_01 PASS helper extracted, both functions present
py_edge_01 PASS first_element + safe_divide both fixed

Rust (Qwen3.5-9B) — 2/5 tasks passing (40%)

Type system friction is the main limiter. No fine-tuning on either language.

Before kernel (same models, free-form): 0% CRP compliance. After kernel: 84% CRP compliance on Qwen3.5-9B.

Test plan

  • cargo test (114 tests)
  • cinto index on real workspace
  • cinto batch --kernel --dry-run validates tasks
  • F5 toggle in TUI, stage indicator visible
  • cinto setup → Esc saves config
  • cinto use-preset lm-studio-small --model <name> writes config
  • Python baseline: 4/5 end-to-end test pass rate
  • Rust baseline: 2/5 end-to-end test pass rate

🤖 Generated with Claude Code

Your Name and others added 30 commits May 15, 2026 15:38
- Added effort-conditional CRP template resolution (minimal, standard, thorough)
- Implemented CRP-aware context compression in session history
- Added `cinto batch` command for headless synthetic dataset generation
- Integrated automated workspace cleanup (`git reset --hard`) for batch tasks
- Added optional LLM-as-a-judge trace evaluator for batch trace generation
- Added granular metrics to BatchResult (tokens, duration, pass rates)
- Added 'fixture_dir' support for 100% isolated and reproducible temp workspaces
- Added 'validation_command' for dynamic semantic testing (e.g. cargo test)
- Added '--dry-run' flag to validate tasks without calling the model
- Included full run metadata (cinto version, model parameters, timestamps) in JSON
- Added `cinto init` command to extract default CRP templates to active workspace
- Created baseline v1 evaluation dataset with 5 stratified Rust tasks
- Configured dynamic `cargo test` validation commands for baseline tasks
- Prevented catastrophic security flaw where headless batch auto-approved all model shell commands on the host by default
- Added `--dangerously-auto-approve` safety flag requirement for batch runs
- Added `Dockerfile.eval` to facilitate safe isolated model evaluation environments
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Introduces src/kernel/ with the first kernel subsystem: index_repo.
Walks the workspace, builds project_map.json (file tree with metadata)
and symbol_index.json (per-file symbol list) in .cinto/.

Symbol extraction covers Rust, Python, JS/TS, and Go via line-level
prefix matching — no new dependencies. The `cinto index` subcommand
exposes this from the CLI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds src/kernel/search.rs with hard caps enforced at every level:
- 15 results total (hard cap 20), 3 per file, 2 context lines (max 4)
- 3000-char output ceiling with explicit truncation signal
- gap indicator (...) when context lines are non-adjacent

Exposes `cinto search <query> [--glob *.rs]` for manual testing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
read_range:   line window (1-indexed, inclusive), capped at 150 lines
read_around:  up to 3 occurrences of a query per file, ±5 lines default
              (max ±20), 4000-char output ceiling
list_symbols: live symbol extraction from disk — always fresh, no content

Exposes read-range, read-around, list-symbols CLI commands for testing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ContextPackBuilder assembles a bounded prompt blob for each worker call:
- [TASK] always included, not counted against code budget
- [REPO MAP] compact file listing from index, hinted files first
- [SYMBOLS] per hinted file — from index if fresh, live extraction fallback
- [SEARCH] scoped search results for each search term hint
- [budget: X / Y chars] footer with truncation signal

Default budget: 16000 chars (~4192 tokens). Configurable via with_budget().
Exposes `cinto pack` CLI for manual inspection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Worker loop drives a 5-stage pipeline for bugfix/feature tasks:
  INTERPRET → LOCATE → HYPOTHESIZE → PATCH → REPORT

Each stage gets a slot-specific context pack (no stage sees more than
it needs). Index I/O is spawned concurrently with model inference via
tokio::spawn_blocking — hiding disk reads behind model call latency.

Typed StageOutput flows between stages. WorkerEvent stream exposes
stage progress, pack budget, retries, and final result for the TUI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
F5 or /kernel toggles between Core (conversational) and Kernel (pipeline).

In kernel mode, user input spawns WorkerLoop::run_bugfix instead of the
agent session. WorkerEvents are drained alongside TurnEvents each tick —
stage name shows in the phase indicator as "kernel · {stage}", per-stage
context pack budget appears in the status bar, retries and failures
surface as transcript items.

is_busy() covers both task types so mode cannot be switched mid-run.
KernelStage(String) added to StreamPhase; render.rs exhaustiveness updated.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
30 unit tests across all deterministic kernel syscalls:
- index: file discovery, skip rules, symbol extraction (Rust/Python/JS)
- search: match caps, glob restriction, char limit (fixed per-line check)
- read: range cap, path traversal rejection, around occurrences cap
- context_pack: budget enforcement, truncation flag, task always present

Also adds kernel batch runner (cinto batch --kernel):
- Runs WorkerLoop pipeline per task, pre-indexes fixture workspace
- Records per-stage CRP validity, retry count, duration
- Outputs JSONL with stages_completed/stages_attempted/workflow_succeeded
- Safe to run without --dangerously-auto-approve (no shell execution)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lm-studio-small preset targets sub-4B models (Gemma 4 E2B, Phi, etc.):
- format: openai-tools (not harmony — LM Studio applies native chat template)
- context_window: 8192, max_tokens: 1024, temperature: 0.1
- thinking_effort: none, stop: [] (no Harmony stop tokens)
- crp_retry_budget: 2, tighter compression settings

WorkerLoop gains with_budget(chars) builder method so the pack budget
flows from CLI → batch runner → worker → each ContextPackBuilder call.

cinto batch --kernel now accepts --kernel-budget <chars> to reduce context
pressure for small models. Recommended for gemma-4-e2b: 6000.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
cinto presets          — lists all built-in presets with descriptions
cinto use-preset <name> [--model <name>]
                       — writes preset to saved config without TUI

Fixes setup-doesn't-persist UX: users can now apply any preset from
the CLI with a single command. --model override lets you apply a preset
but swap in a specific model name (e.g. use lm-studio-small but point
at a different model loaded in LM Studio).

Also fixes cinto batch --kernel failing when output directory doesn't
exist — now creates parent dirs automatically.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Small models like Gemma 4 E2B output plain prose instead of CRP slots.
Previously this caused the entire pipeline to fail at stage 1.

parse_stage_output now tries CRP first; if that fails and the response
is non-empty, it salvages a StageOutput from plain text:
- final_response = full response text
- search_terms extracted from prose (for interpret stage)
- file paths scanned from prose (for locate stage)
- approach = full text (for hypothesize/patch stages)

StageCompleted now carries crp_valid: bool so batch metrics and the TUI
can distinguish structured CRP output from plain-text fallback.
Status bar shows "kernel: interpret ~ (plain text)" vs "✓".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Small models ignore format instructions but respond to concrete examples.
Each stage now includes a minimal before/after example immediately after
the CRP brief, showing the exact slot tags the model must emit:

- interpret: TASK_INTERPRETATION + FINAL_RESPONSE
- locate:    RELEVANT_FILES + FINAL_RESPONSE
- hypothesize: PROPOSED_APPROACH + FINAL_RESPONSE
- patch:     FILE_EDITS + FINAL_RESPONSE
- report:    FINAL_RESPONSE only

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Gemma was emitting `- src/lib.rs` literally from the bullet in the example.
Bare paths match what the model naturally produces.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Allows overriding context window and max output tokens when applying a
preset, without having to manually edit the config file afterwards.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Qwen3, DeepSeek-R1 and similar reasoning models prefix their response
with a <think>...</think> chain. With constrained max_tokens this block
can exhaust the budget before any actual output is produced, causing
empty-response failures.

strip_think_blocks() finds the last </think> tag and parses everything
after it, so the actual answer is always visible to the CRP parser.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Your Name and others added 13 commits May 15, 2026 15:38
Each successful kernel stage emits WorkerEvent::StageTrace carrying the
full (system_prompt, user_message, model_response, crp_valid) triple.

cinto batch --kernel --save-traces <dir> writes these to a JSONL file
in the given directory, one record per stage per task. Records include
task_id, model, stage, workflow_succeeded, and timestamp so they can be
filtered downstream (e.g. keep only crp_valid=true + workflow_succeeded=true).

The TUI silently discards StageTrace events — they are batch-only.

Also replaces goal.md with the full strategic roadmap:
Kernel → Dataset → CintoLM fine-tune → Constrained Decoding → Architecture.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When locate exhausts its retry budget or returns empty, the pipeline
previously propagated the error and killed everything downstream.

Now: locate failure emits StageSkipped and returns a default StageOutput.
Hypothesize already passes interpret's search_terms to ContextPackBuilder,
which runs ripgrep searches to find relevant code — so the pipeline
continues with search-based context even without explicit file identification.

StageSkipped is surfaced as a transcript item in the TUI and recorded in
the batch errors list for analysis.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
src/kernel/patch.rs:
- parse_edit_directives: handles <EDIT> CRP blocks, @@ terse format
- EditMode: Replace, ReplaceFunction(name), Prepend, Append
- apply_directive: writes files with path traversal validation
- replace_function_body: brace-depth function replacement
- preview: compact before/after summary for approval prompt

worker.rs: after PATCH stage, emit PatchApprovalRequested per directive,
await user response, apply approved edits, feed summary into REPORT stage.

ui.rs: PatchApprovalRequested reuses existing PendingToolApproval flow
(y/Enter approve, n/Esc reject).

batch.rs: concurrent event handler task auto-approves patches and records
all other events — fixes the deadlock where response_rx.await would block
the workflow while events were collected post-hoc.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ModelConfig gains no_think: bool and no_think_prefix: String. When
no_think is true, Config::apply_no_think_prefix() prepends the prefix
to the system prompt, instructing Qwen3-class models to skip their
<think> chain entirely — more efficient than stripping blocks post-hoc.

New preset: qwen3 — openai-tools, 32K context, /no_think prefix,
crp_retry_budget 2. Works for qwen3.5-9b and compatible models.

Also adds scripts/traces_to_training_jsonl.py and gitignores the
generated eval/results/, eval/traces/, eval/training/ directories.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously the setup screen only persisted on explicit "Save & Start".
Now any exit (Esc, Tab, F2) saves the current config, which is the
expected behaviour — changes shouldn't silently disappear.

Context window auto-detection still only runs on "Save & Start" since
it requires a live server connection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
After a successful workflow where files were changed, the kernel batch
runner now executes the task's validation_command (e.g. cargo test) in
the fixture workspace and records validation_passed + validation_output
in the JSONL result.

This gives the real metric: does the kernel-generated patch compile and
pass tests — not just whether the pipeline produced structured output.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The patch stage previously only sent symbol tables for located files.
A model writing a replace-mode edit needs the actual file content or it
writes only the changed piece, discarding everything else.

Now passes explicit ranges (lines 1–300) for each located file so the
model can see what it is replacing before generating FILE_EDITS.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
patch.rs: replace_function_body now detects Python by checking if the
declaration line ends with ':'. Uses indentation-based scoping instead
of brace-depth for Python files.

5 Python baseline tasks (unittest, python3 -m unittest test_lib):
- py_bugfix_01: off-by-one in find_max range()
- py_bugfix_02: ZeroDivisionError on empty list in calculate_average
- py_feature_01: add get_username/get_email getters to User class
- py_refactor_01: extract calculate_average helper from process_data
- py_edge_01: first_element and safe_divide return None on bad input

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When the model generates replace_function:name for a function that
doesn't exist yet (e.g. a newly extracted helper), the patch applier
now appends it to the file instead of silently failing.

This fixes the refactor pattern: the model replaces the original
function and adds the extracted helper — both edits succeed even
though the helper is new.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 1-300 range added full file content to the PATCH stage context pack,
which pushed the system prompt over 4096 tokens on models loaded with
a small context window in LM Studio (n_keep >= n_ctx error).

80 lines covers the vast majority of small task files while keeping the
context pack well within 4096-token budgets.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 80-line cap was a band-aid for LM Studio loading models with a
smaller n_ctx than configured (e.g. 4096 instead of 32768).

Now run_kernel probes the server's actual context window at startup.
If it's smaller than config.model.context_window:
  - Prints a warning with the exact LM Studio setting to fix
  - Computes a safe budget from the real window (70% at 4 chars/token)
  - Uses that budget for all tasks in the run

patch stage range restored to 1-150 (from 80) since budget now
correctly governs total pack size regardless of model context.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
memory.rs: save_decision writes HYPOTHESIZE approach to
.cinto/memory/decisions.md with timestamp. load_recent_decisions
returns last 2K chars for injection into future HYPOTHESIZE context
packs — approaches from past runs inform new ones.

worker.rs: loads past decisions before HYPOTHESIZE stage and saves
the approach after. Completes the v0.2 spec memory requirement.

Also fixes the PATCH one-shot example: shows replace mode with the
COMPLETE class body for adding methods, alongside replace_function
for targeted single-function edits. Addresses the py_feature_01
pattern where getters were appended outside the class.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
@Jshebb
Jshebb force-pushed the feat/cinto-kernel-architecture branch from b7e538f to c20d474 Compare May 15, 2026 18:41
@Jshebb
Jshebb merged commit 5b09ca9 into master May 15, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant