Repository navigation
feat: Cinto Kernel — pipeline runtime for local AI - #1
Merged
Merged
Conversation
- Added effort-conditional CRP template resolution (minimal, standard, thorough) - Implemented CRP-aware context compression in session history - Added `cinto batch` command for headless synthetic dataset generation - Integrated automated workspace cleanup (`git reset --hard`) for batch tasks - Added optional LLM-as-a-judge trace evaluator for batch trace generation
- Added granular metrics to BatchResult (tokens, duration, pass rates) - Added 'fixture_dir' support for 100% isolated and reproducible temp workspaces - Added 'validation_command' for dynamic semantic testing (e.g. cargo test) - Added '--dry-run' flag to validate tasks without calling the model - Included full run metadata (cinto version, model parameters, timestamps) in JSON
- Added `cinto init` command to extract default CRP templates to active workspace - Created baseline v1 evaluation dataset with 5 stratified Rust tasks - Configured dynamic `cargo test` validation commands for baseline tasks
- Prevented catastrophic security flaw where headless batch auto-approved all model shell commands on the host by default - Added `--dangerously-auto-approve` safety flag requirement for batch runs - Added `Dockerfile.eval` to facilitate safe isolated model evaluation environments
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Introduces src/kernel/ with the first kernel subsystem: index_repo. Walks the workspace, builds project_map.json (file tree with metadata) and symbol_index.json (per-file symbol list) in .cinto/. Symbol extraction covers Rust, Python, JS/TS, and Go via line-level prefix matching — no new dependencies. The `cinto index` subcommand exposes this from the CLI. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds src/kernel/search.rs with hard caps enforced at every level: - 15 results total (hard cap 20), 3 per file, 2 context lines (max 4) - 3000-char output ceiling with explicit truncation signal - gap indicator (...) when context lines are non-adjacent Exposes `cinto search <query> [--glob *.rs]` for manual testing. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
read_range: line window (1-indexed, inclusive), capped at 150 lines
read_around: up to 3 occurrences of a query per file, ±5 lines default
(max ±20), 4000-char output ceiling
list_symbols: live symbol extraction from disk — always fresh, no content
Exposes read-range, read-around, list-symbols CLI commands for testing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ContextPackBuilder assembles a bounded prompt blob for each worker call: - [TASK] always included, not counted against code budget - [REPO MAP] compact file listing from index, hinted files first - [SYMBOLS] per hinted file — from index if fresh, live extraction fallback - [SEARCH] scoped search results for each search term hint - [budget: X / Y chars] footer with truncation signal Default budget: 16000 chars (~4192 tokens). Configurable via with_budget(). Exposes `cinto pack` CLI for manual inspection. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Worker loop drives a 5-stage pipeline for bugfix/feature tasks: INTERPRET → LOCATE → HYPOTHESIZE → PATCH → REPORT Each stage gets a slot-specific context pack (no stage sees more than it needs). Index I/O is spawned concurrently with model inference via tokio::spawn_blocking — hiding disk reads behind model call latency. Typed StageOutput flows between stages. WorkerEvent stream exposes stage progress, pack budget, retries, and final result for the TUI. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
F5 or /kernel toggles between Core (conversational) and Kernel (pipeline).
In kernel mode, user input spawns WorkerLoop::run_bugfix instead of the
agent session. WorkerEvents are drained alongside TurnEvents each tick —
stage name shows in the phase indicator as "kernel · {stage}", per-stage
context pack budget appears in the status bar, retries and failures
surface as transcript items.
is_busy() covers both task types so mode cannot be switched mid-run.
KernelStage(String) added to StreamPhase; render.rs exhaustiveness updated.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
30 unit tests across all deterministic kernel syscalls: - index: file discovery, skip rules, symbol extraction (Rust/Python/JS) - search: match caps, glob restriction, char limit (fixed per-line check) - read: range cap, path traversal rejection, around occurrences cap - context_pack: budget enforcement, truncation flag, task always present Also adds kernel batch runner (cinto batch --kernel): - Runs WorkerLoop pipeline per task, pre-indexes fixture workspace - Records per-stage CRP validity, retry count, duration - Outputs JSONL with stages_completed/stages_attempted/workflow_succeeded - Safe to run without --dangerously-auto-approve (no shell execution) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
lm-studio-small preset targets sub-4B models (Gemma 4 E2B, Phi, etc.): - format: openai-tools (not harmony — LM Studio applies native chat template) - context_window: 8192, max_tokens: 1024, temperature: 0.1 - thinking_effort: none, stop: [] (no Harmony stop tokens) - crp_retry_budget: 2, tighter compression settings WorkerLoop gains with_budget(chars) builder method so the pack budget flows from CLI → batch runner → worker → each ContextPackBuilder call. cinto batch --kernel now accepts --kernel-budget <chars> to reduce context pressure for small models. Recommended for gemma-4-e2b: 6000. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
cinto presets — lists all built-in presets with descriptions
cinto use-preset <name> [--model <name>]
— writes preset to saved config without TUI
Fixes setup-doesn't-persist UX: users can now apply any preset from
the CLI with a single command. --model override lets you apply a preset
but swap in a specific model name (e.g. use lm-studio-small but point
at a different model loaded in LM Studio).
Also fixes cinto batch --kernel failing when output directory doesn't
exist — now creates parent dirs automatically.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Small models like Gemma 4 E2B output plain prose instead of CRP slots. Previously this caused the entire pipeline to fail at stage 1. parse_stage_output now tries CRP first; if that fails and the response is non-empty, it salvages a StageOutput from plain text: - final_response = full response text - search_terms extracted from prose (for interpret stage) - file paths scanned from prose (for locate stage) - approach = full text (for hypothesize/patch stages) StageCompleted now carries crp_valid: bool so batch metrics and the TUI can distinguish structured CRP output from plain-text fallback. Status bar shows "kernel: interpret ~ (plain text)" vs "✓". Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Small models ignore format instructions but respond to concrete examples. Each stage now includes a minimal before/after example immediately after the CRP brief, showing the exact slot tags the model must emit: - interpret: TASK_INTERPRETATION + FINAL_RESPONSE - locate: RELEVANT_FILES + FINAL_RESPONSE - hypothesize: PROPOSED_APPROACH + FINAL_RESPONSE - patch: FILE_EDITS + FINAL_RESPONSE - report: FINAL_RESPONSE only Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Gemma was emitting `- src/lib.rs` literally from the bullet in the example. Bare paths match what the model naturally produces. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Allows overriding context window and max output tokens when applying a preset, without having to manually edit the config file afterwards. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Qwen3, DeepSeek-R1 and similar reasoning models prefix their response with a <think>...</think> chain. With constrained max_tokens this block can exhaust the budget before any actual output is produced, causing empty-response failures. strip_think_blocks() finds the last </think> tag and parses everything after it, so the actual answer is always visible to the CRP parser. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Each successful kernel stage emits WorkerEvent::StageTrace carrying the full (system_prompt, user_message, model_response, crp_valid) triple. cinto batch --kernel --save-traces <dir> writes these to a JSONL file in the given directory, one record per stage per task. Records include task_id, model, stage, workflow_succeeded, and timestamp so they can be filtered downstream (e.g. keep only crp_valid=true + workflow_succeeded=true). The TUI silently discards StageTrace events — they are batch-only. Also replaces goal.md with the full strategic roadmap: Kernel → Dataset → CintoLM fine-tune → Constrained Decoding → Architecture. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When locate exhausts its retry budget or returns empty, the pipeline previously propagated the error and killed everything downstream. Now: locate failure emits StageSkipped and returns a default StageOutput. Hypothesize already passes interpret's search_terms to ContextPackBuilder, which runs ripgrep searches to find relevant code — so the pipeline continues with search-based context even without explicit file identification. StageSkipped is surfaced as a transcript item in the TUI and recorded in the batch errors list for analysis. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
src/kernel/patch.rs: - parse_edit_directives: handles <EDIT> CRP blocks, @@ terse format - EditMode: Replace, ReplaceFunction(name), Prepend, Append - apply_directive: writes files with path traversal validation - replace_function_body: brace-depth function replacement - preview: compact before/after summary for approval prompt worker.rs: after PATCH stage, emit PatchApprovalRequested per directive, await user response, apply approved edits, feed summary into REPORT stage. ui.rs: PatchApprovalRequested reuses existing PendingToolApproval flow (y/Enter approve, n/Esc reject). batch.rs: concurrent event handler task auto-approves patches and records all other events — fixes the deadlock where response_rx.await would block the workflow while events were collected post-hoc. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ModelConfig gains no_think: bool and no_think_prefix: String. When no_think is true, Config::apply_no_think_prefix() prepends the prefix to the system prompt, instructing Qwen3-class models to skip their <think> chain entirely — more efficient than stripping blocks post-hoc. New preset: qwen3 — openai-tools, 32K context, /no_think prefix, crp_retry_budget 2. Works for qwen3.5-9b and compatible models. Also adds scripts/traces_to_training_jsonl.py and gitignores the generated eval/results/, eval/traces/, eval/training/ directories. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously the setup screen only persisted on explicit "Save & Start". Now any exit (Esc, Tab, F2) saves the current config, which is the expected behaviour — changes shouldn't silently disappear. Context window auto-detection still only runs on "Save & Start" since it requires a live server connection. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
After a successful workflow where files were changed, the kernel batch runner now executes the task's validation_command (e.g. cargo test) in the fixture workspace and records validation_passed + validation_output in the JSONL result. This gives the real metric: does the kernel-generated patch compile and pass tests — not just whether the pipeline produced structured output. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The patch stage previously only sent symbol tables for located files. A model writing a replace-mode edit needs the actual file content or it writes only the changed piece, discarding everything else. Now passes explicit ranges (lines 1–300) for each located file so the model can see what it is replacing before generating FILE_EDITS. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
patch.rs: replace_function_body now detects Python by checking if the declaration line ends with ':'. Uses indentation-based scoping instead of brace-depth for Python files. 5 Python baseline tasks (unittest, python3 -m unittest test_lib): - py_bugfix_01: off-by-one in find_max range() - py_bugfix_02: ZeroDivisionError on empty list in calculate_average - py_feature_01: add get_username/get_email getters to User class - py_refactor_01: extract calculate_average helper from process_data - py_edge_01: first_element and safe_divide return None on bad input Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When the model generates replace_function:name for a function that doesn't exist yet (e.g. a newly extracted helper), the patch applier now appends it to the file instead of silently failing. This fixes the refactor pattern: the model replaces the original function and adds the extracted helper — both edits succeed even though the helper is new. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 1-300 range added full file content to the PATCH stage context pack, which pushed the system prompt over 4096 tokens on models loaded with a small context window in LM Studio (n_keep >= n_ctx error). 80 lines covers the vast majority of small task files while keeping the context pack well within 4096-token budgets. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The 80-line cap was a band-aid for LM Studio loading models with a smaller n_ctx than configured (e.g. 4096 instead of 32768). Now run_kernel probes the server's actual context window at startup. If it's smaller than config.model.context_window: - Prints a warning with the exact LM Studio setting to fix - Computes a safe budget from the real window (70% at 4 chars/token) - Uses that budget for all tasks in the run patch stage range restored to 1-150 (from 80) since budget now correctly governs total pack size regardless of model context. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
memory.rs: save_decision writes HYPOTHESIZE approach to .cinto/memory/decisions.md with timestamp. load_recent_decisions returns last 2K chars for injection into future HYPOTHESIZE context packs — approaches from past runs inform new ones. worker.rs: loads past decisions before HYPOTHESIZE stage and saves the approach after. Completes the v0.2 spec memory requirement. Also fixes the PATCH one-shot example: shows replace mode with the COMPLETE class body for adding methods, alongside replace_function for targeted single-function edits. Addresses the py_feature_01 pattern where getters were appended outside the class. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Jshebb
force-pushed
the
feat/cinto-kernel-architecture
branch
from
May 15, 2026 18:41
b7e538f to
c20d474
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds Cinto Kernel — a pipeline runtime that turns small local models into reliable coding agents by giving each stage one bounded job instead of asking the model to do everything at once.
What's in this PR
index_repo,search,read_range,read_around,list_symbols,context_pack— all with hard output caps so models aren't overwhelmed<EDIT>CRP blocks and@@terse format;replace_functionfalls back to append for new functions; indentation-based scoping for Python.cinto/memory/decisions.mdwritten after each HYPOTHESIZE stage; loaded into future runs for context across sessions/kerneltoggles between Core (conversational) and Kernel (pipeline) modeslm-studio-small+qwen3presets,no_thinkconfig field, think-block stripping, plain-text fallback, per-stage one-shot CRP examplescinto batch --kernel, context window auto-detection,cinto use-preset,cinto presetsEval results (baseline tasks, no fine-tuning)
Python (Qwen3.5-9B) — 4/5 tasks passing (80%)
find_maxoff-by-one fixedcalculate_averageempty list handledfirst_element+safe_divideboth fixedRust (Qwen3.5-9B) — 2/5 tasks passing (40%)
Type system friction is the main limiter. No fine-tuning on either language.
Before kernel (same models, free-form): 0% CRP compliance. After kernel: 84% CRP compliance on Qwen3.5-9B.
Test plan
cargo test(114 tests)cinto indexon real workspacecinto batch --kernel --dry-runvalidates taskscinto setup→ Esc saves configcinto use-preset lm-studio-small --model <name>writes config🤖 Generated with Claude Code