Repository navigation
Expand file tree
/
Copy pathCodeCompass.html
More file actions
318 lines (298 loc) · 30.4 KB
/
Copy pathCodeCompass.html
File metadata and controls
318 lines (298 loc) · 30.4 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>CodeCompass</title>
<style>
:root{--ink:#1c2530;--muted:#5b6b7b;--line:#e2e8f0;--bg:#ffffff;--soft:#f6f8fb;--accent:#2563eb;--good:#15803d;--warn:#b45309;--code:#0f172a;--codebg:#f1f5f9}
*{box-sizing:border-box}
body{margin:0;background:var(--bg);color:var(--ink);font:15px/1.6 -apple-system,Segoe UI,Roboto,Helvetica,Arial,sans-serif}
.wrap{max-width:900px;margin:0 auto;padding:40px 24px 80px}
h1{font-size:30px;margin:0 0 2px}
.tag{color:var(--muted);font-size:16px;margin:0 0 28px}
h2{font-size:20px;margin:36px 0 10px;padding-bottom:6px;border-bottom:1px solid var(--line)}
h3{font-size:16px;margin:22px 0 6px}
p{margin:8px 0}
ul{margin:8px 0 8px 22px;padding:0}
li{margin:4px 0}
code{background:var(--codebg);color:var(--code);padding:1px 6px;border-radius:5px;font:13px/1.5 ui-monospace,Consolas,monospace}
pre{background:var(--codebg);color:var(--code);padding:12px 14px;border-radius:8px;overflow:auto;font:13px/1.55 ui-monospace,Consolas,monospace}
table{width:100%;border-collapse:collapse;margin:12px 0;font-size:14px}
th,td{padding:7px 10px;border-bottom:1px solid var(--line);text-align:right;white-space:nowrap}
th{background:var(--soft);color:var(--muted);font-weight:600;font-size:12px;text-transform:uppercase;letter-spacing:.03em}
td:first-child,th:first-child,td:nth-child(2),th:nth-child(2){text-align:left}
.num{font-variant-numeric:tabular-nums}
.card{background:var(--soft);border:1px solid var(--line);border-radius:10px;padding:14px 16px;margin:14px 0}
.pill{display:inline-block;background:#eef2ff;color:var(--accent);border-radius:999px;padding:2px 10px;font-size:12px;font-weight:600;margin:0 4px 4px 0}
.yes{color:var(--good);font-weight:600}.no{color:var(--warn);font-weight:600}
.muted{color:var(--muted)}
footer{margin-top:40px;color:var(--muted);font-size:13px;border-top:1px solid var(--line);padding-top:14px}
</style>
</head>
<body><div class="wrap">
<h1>CodeCompass</h1>
<p class="tag">Fast, fully-local code search & navigation for large codebases — built to cut the tokens an AI coding agent spends finding code.</p>
<div class="card">
<b>In one line:</b> CodeCompass indexes your repo on your machine and gives Claude Code (or your terminal) precise
<code>file:line:col</code> answers — text search, go-to-definition, and find-references — instead of the agent grepping
blindly and reading whole files. Nothing leaves your machine. No GPU. No cloud.
</div>
<h2>What it is & why</h2>
<p>An AI coding agent's default moves — grep for a string, then read entire files to understand a match — are cheap on a
small repo and ruinously token-hungry on a big one. CodeCompass returns just the lines that matter, so the agent spends its
context on thinking, not on scrolling files it didn't need.</p>
<p>It's three complementary layers, cheapest first:</p>
<ul>
<li><b>Lexical</b> — a trigram index for instant literal/substring search across the whole repo.</li>
<li><b>Symbolic</b> — tree-sitter parses every file into symbols (classes, methods, functions…) for go-to-definition.</li>
<li><b>Semantic</b> — real find-references for <b>C#</b> (Roslyn): it resolves the actual symbol. Every other language,
<b>C and C++ included</b>, gets a fast <b>name search</b> for references: whole-word uses in code, never the definition
itself, and for C and C++ never inside a comment or string.</li>
</ul>
<h3>Minimal token footprint</h3>
<p>A pain point with some in-house tools is that they register a pile of MCP tools/skills that eat context every session.
CodeCompass exposes a deliberately small <b>7 tools</b> with terse descriptions — that's the entire standing per-session cost.</p>
<p><span class="pill">search_code</span><span class="pill">find_definition</span><span class="pill">find_references</span><span class="pill">find_callees</span><span class="pill">search_symbols</span><span class="pill">reindex</span><span class="pill">manage_links</span></p>
<h2>What it can & can't do</h2>
<table>
<thead><tr><th>Capability</th><th style="text-align:left">Status</th></tr></thead>
<tbody>
<tr><td>Literal / substring search (all files)</td><td style="text-align:left"><span class="yes">Yes</span></td></tr>
<tr><td>Go-to-definition & symbol search</td><td style="text-align:left"><span class="yes">Yes</span> — C#, C, C++, Python, JS, TS/TSX, Go, Rust, TRACE32 PRACTICE (.cmm)</td></tr>
<tr><td>Find-references</td><td style="text-align:left"><span class="yes">Yes</span> — semantic for C#; by name elsewhere (C/C++: comments, strings and definitions excluded)</td></tr>
<tr><td>Auto re-index on file changes (get-latest, edits)</td><td style="text-align:left"><span class="yes">Yes</span> — debounced, content-hash verified, ignores build output</td></tr>
<tr><td>Runs fully local, no GPU, no cloud</td><td style="text-align:left"><span class="yes">Yes</span></td></tr>
<tr><td>Dozens-of-GB repos without exhausting RAM</td><td style="text-align:left"><span class="yes">Yes</span> — indexes are memory-mapped on disk</td></tr>
<tr><td>MATLAB / other unlisted languages</td><td style="text-align:left"><span class="no">Lexical only</span> (text search works; no symbols)</td></tr>
<tr><td>Very large files</td><td style="text-align:left"><span class="no">Capped</span> — >5 MB skipped from the index; >1 MB skipped from symbols but still text-searchable (both tunable)</td></tr>
<tr><td>C/C++ find-references</td><td style="text-align:left"><span class="no">By name</span>, not compiled: fast and complete on any repo, but same-named symbols are listed together</td></tr>
<tr><td>Semantic "meaning" / embedding search</td><td style="text-align:left"><span class="no">No</span> (deliberately — needs a model; weaker for real code nav)</td></tr>
</tbody>
</table>
<h2>Install</h2>
<p>Requires the .NET 8 SDK to <i>build</i>; the produced binaries are self-contained (no .NET needed to run them). Windows x64.</p>
<p class="muted"><b>Windows on ARM (Snapdragon):</b> no separate ARM build is needed — the x64 binaries run on Windows 11 on ARM through its built-in x64 emulation (the whole process, native dependencies included, runs emulated). Searches stay effectively instant; only the first index build runs somewhat slower than on native x64.</p>
<pre>pwsh ./build-plugin.ps1 # publishes self-contained binaries into plugin/bin</pre>
<p>Then load it into Claude Code (there is no <code>/plugin add</code> command):</p>
<pre>claude --plugin-dir "<path>/plugin" # one session
/plugin marketplace add "<path>/plugin" # persistent: register the folder…
/plugin install codecompass@codecompass # …then install (run inside Claude Code)</pre>
<p>The prebuilt release zip unzips to a folder you point at the same way (it ships a self-marketplace). In a session, run <code>/mcp</code> to confirm the <b>codecompass</b> server is connected.</p>
<h2>Using it from Claude Code</h2>
<p>Once installed you don't do anything special — just ask Claude to work in your repo. It gets the CodeCompass tools, and a
hook <b>redirects <code>Grep</code> to CodeCompass</b> so the agent uses the index instead of scanning files (<code>Glob</code> stays available — CodeCompass searches contents and symbols, not file names). (Set
<code>CODECOMPASS_ENFORCE=0</code> to allow grep again.) The index builds automatically on first use for normal-sized workspaces
and stays fresh as files change.</p>
<p><b>Large repos:</b> a very large workspace isn't auto-indexed inside a tool call (that could stall). CodeCompass tells you to
build it once from a terminal — see below — then it serves and keeps it fresh.</p>
<h2>Using it from the command line</h2>
<pre>codecompass index <path> build the index (shows progress + ETA)
codecompass update <path> incremental reindex of changes
codecompass watch <path> auto-reindex on file changes
codecompass search <path> <query> literal text search -> file:line:col
codecompass def <path> <name> go-to-definition
codecompass refs <path> <name> references (semantic C#; by name elsewhere)
codecompass callees <path> <name> in-repo methods a C# method calls (C# only)
codecompass symbols <path> <substr> symbol-name search
codecompass survey <path> report what the size caps skip + suggest config
codecompass logs show the log folder and files
codecompass version print the build version (1.0.<commit-count>+<short-sha>)
codecompass symstats <path> profile symbol-file sizes + parse cost per language
codecompass parsebench tree-sitter parse-time vs size sweep (synthetic)
codecompass statusline [--wrap "..."] Claude Code status-line segment (index state)
codecompass doctor <path> diagnose a repo's index (health + metadata)
codecompass cache [list|gc|clear <path>] inspect/manage the per-user index cache
codecompass report <path> [--no-logs] zip diagnostics + logs for a bug report (never source)</pre>
<h2>Linked roots (searching more than one repo)</h2>
<p>When the code you work on spans more than one place — a shared library checked out elsewhere, a sibling
repo, or a third-party drop on another drive that can't be nested under your project — <b>linked roots</b> let a
project index and search those external directories alongside its own, as one federated result set. Each linked root
keeps its <b>own</b> index (keyed by its absolute path, so a root two projects both link is indexed once and reused),
is live-watched for edits like the main repo, and <code>find_references</code>/<code>find_callees</code> resolve
<b>across</b> the boundary (C# via Roslyn; name matches search every root). Links are stored machine-local (absolute paths), not in
the committed <code>.codecompass.json</code>.</p>
<p>Manage them from <b>either</b> side — they share one implementation:</p>
<pre>From the agent (MCP tool): manage_links action: list | add | remove | focus (path: for add/remove/focus)
From a terminal (CLI): codecompass link add|remove|list <path> [project-dir]</pre>
<p><b>Focus:</b> linked several big repos but working in one? <code>manage_links action=focus path="<repo>"</code>
scopes the session's searches to it (a folder name, path fragment or absolute path; comma-separated for several; no
path clears it). Every scoped answer says what was excluded, and a focused root with no index yet is called out, so a
narrowed search is never mistaken for "not found". Focus filters what's shown, not what's resolved.</p>
<p><b>Only the MCP server searches across linked roots</b> (it's the long-lived process that holds each root's index
open and watches it); the CLI always searches the single root you point it at. So: set federation up with
<code>manage_links</code> or <code>link add</code>, then let the MCP tools do the multi-repo searching.
<code>codecompass doctor <project></code> lists every linked root, whether it exists and is indexed, and how many
other projects share it.</p>
<h2>Tuning for huge or generated trees</h2>
<p>Indexing is robust on ordinary source at scale. The one hazard is <i>large machine-generated files</i>. Tree-sitter parse
cost is <b>linear in file size</b>, but the constant factor varies ~70× by content (measured with <code>parsebench</code>):
ordinary code parses at ~1.8 MB/s, while the worst case measured — deeply nested C++ templates — runs at
~0.1 MB/s (a 1 MB file ≈ 12 s). So a big degenerate file can take tens of seconds even though it grows
linearly. That is bounded by default — <b>symbol extraction is skipped above <code>CODECOMPASS_MAX_SYMBOL_MB</code>
(1 MB)</b>, and those files are still fully trigram-indexed, so text search stays complete. In practice it never touches
hand-written code: across the test corpora, every source file over 1 MB was machine-generated. Files skipped for a cap are
counted and logged, so the coverage gap is visible rather than silent.</p>
<p>Two diagnostics quantify this on your own repo: <code>codecompass symstats <path></code> reports the per-language file-size
distribution (from <code>stat</code> only — free, so it is cheap on a huge repo) and runs tree-sitter, ignoring the cap,
only on the largest ~100 files per language (sequential, timeout-guarded — it can neither hang nor crawl; add
<code>--full</code> to parse everything) to show whether the big files yield symbols and where parse cost spikes;
<code>codecompass parsebench</code> runs the synthetic size sweep above. A good <code>maxSymbolMb</code> sits above your
largest real symbol-bearing file and below the size where parsing gets slow.</p>
<table>
<thead><tr><th>Environment variable</th><th style="text-align:left">Effect</th></tr></thead>
<tbody>
<tr><td><code>CODECOMPASS_MAX_SYMBOL_MB</code></td><td style="text-align:left">Skip tree-sitter symbol extraction above this size (default 1). Raise it for large <i>valid</i> code whose symbols you want — safe from the data-blob crawl, since above 1 MB overwhelmingly numeric/hex files are auto-skipped for symbols by content.</td></tr>
<tr><td><code>CODECOMPASS_MAX_FILE_MB</code></td><td style="text-align:left">Per-file size cap for indexing entirely (default 2000 / 2 GB). Files ≥128 MB are streamed (bounded memory), so a high cap won't blow up RAM; lower it per-repo to skip big generated files. Its cost is read time on a full build.</td></tr>
<tr><td><code>CODECOMPASS_IGNORE</code></td><td style="text-align:left">Comma/semicolon-separated directory names to exclude (e.g. <code>generated,vendor</code>).</td></tr>
<tr><td><code>CODECOMPASS_KEEP</code></td><td style="text-align:left">Directory names to index even though they're skipped by default as build output (e.g. <code>packages</code> in a pnpm/yarn monorepo). Config: <code>keepDirs</code>. A zero result says when such directories were skipped.</td></tr>
<tr><td><code>CODECOMPASS_MAX_AUTO_MB</code></td><td style="text-align:left">Workspaces larger than this (default 100) are left for a one-time CLI build instead of auto-indexing in a tool call.</td></tr>
<tr><td><code>CODECOMPASS_THREADS</code> / <code>CODECOMPASS_SEGMENT_MB</code></td><td style="text-align:left">Indexing parallelism / per-worker build-memory budget.</td></tr>
<tr><td><code>CODECOMPASS_READ_BUDGET_MB</code></td><td style="text-align:left">Cap on total file bytes held in memory at once during a parallel build (default scales to RAM: ~1/16th of available, clamped 256 MB–4 GB). Stops N cores from each loading a multi-GB file at once when <code>MAX_FILE_MB</code> is large; an over-budget file reads solo.</td></tr>
<tr><td><code>CODECOMPASS_ENFORCE=0</code></td><td style="text-align:left">Allow the agent to use Grep again (the hook redirects it to CodeCompass by default).</td></tr>
</tbody>
</table>
<p>Every knob also lives in an optional <b><code>.codecompass.json</code></b> at the repo root, so settings travel with the
repo (precedence: <b>env → config file → default</b>). Fields: <code>maxSymbolMb</code>, <code>maxFileMb</code>,
<code>maxAutoMb</code>, <code>ignore</code> (directory names), <code>threads</code>, <code>segmentMb</code>,
<code>compactSegments</code>, <code>stallWarnSec</code>, <code>readBudgetMb</code>, <code>autoReconcile</code>, <code>statusLine</code>. Run <code>codecompass survey <path></code> first: it reports what
the caps skip, names the largest files, and suggests a change. It never auto-raises a cap;
file size isn't a reliable signal of parse safety, so that call is left to the repo owner.</p>
<h2>Staying fresh & status line</h2>
<p>While Claude is running, a file watcher keeps the index current. Changes made <i>outside</i> a session
— a Perforce/git sync, a branch switch with Claude closed — have no watcher to catch them, so on
startup the server <b>reconciles</b> the loaded index against the current tree (a stat-walk vs. the last
snapshot) in the background, while the old index keeps serving, then swaps. Automatic for local repos within
<code>maxAutoMb</code>; deferred to a manual <code>codecompass update</code> for network shares and huge repos
(a full-tree stat-walk is slow over SMB and the watcher is unreliable there). Force it with
<code>autoReconcile</code> / <code>CODECOMPASS_AUTO_RECONCILE</code> (<code>true</code>/<code>false</code>).</p>
<p>CodeCompass can also show its state in Claude Code's status area. The server publishes state to a tiny per-repo
file; the <code>codecompass statusline</code> command reads it. Add to Claude Code's <code>settings.json</code>:
<code>{ "statusLine": { "type": "command", "command": "codecompass statusline" } }</code>. Since setting any
status-line command <i>replaces</i> Claude's built-in default, the bare command renders a self-contained line so you
don't lose the usual info — <code>Opus 4.8 · ~/CodeCompass · CodeCompass ✓ 48,000 files</code>
(states: <code>… indexing 42%</code>, <code>↻ refreshing…</code>, <code>⚠ not indexed</code>;
the CodeCompass part drops in un-indexed repos). There is a single status-line slot — if you already have your own
line, use <code>--wrap "<cmd>"</code>: it forwards Claude's stdin, prints your command's output, and appends
<i>only</i> the CodeCompass segment (not model/cwd — your command already shows those). Off with
<code>statusLine: false</code> / <code>CODECOMPASS_STATUS_LINE=0</code>.</p>
<h2>Diagnostics & bug reports</h2>
<p><code>codecompass doctor <path></code> is a read-only health check: version, environment, config presence, index
metadata (documents/segments, the version + time that built it), the cache-file listing, and pass/warn checks (builds?
loads cleanly? current version? network path?). <code>codecompass cache</code> lists/GCs/clears the per-user cache by
real repo path. <code>codecompass report <path></code> zips the <code>doctor</code> snapshot + this repo's logs +
its <code>.codecompass.json</code> for a bug report — and <b>nothing from the source tree</b>: it contains file
paths/names and sizes (from logs), never file <i>contents</i> or index contents (<code>--no-logs</code> makes a
paths-free bundle). Safe to run on someone else's proprietary code.</p>
<h2>Limitations</h2>
<p>Worth knowing where the edges are — all are by design and all skips are logged, never silent:</p>
<ul>
<li>Files over <code>MAX_FILE_MB</code> (default 2000 MB / 2 GB) are absent from search entirely.</li>
<li>Files over <code>MAX_SYMBOL_MB</code> (default 1 MB), classified as numeric/hex data above 1 MB when the cap is raised, or streamed (≥128 MB), have no go-to-definition (still text-searchable).</li>
<li>Files <b>≥128 MB are indexed by streaming</b> (bounded memory — a 2 GB file indexes even on 16 GB), removing the old ~1 GB single-string ceiling. A search into one uses a per-file <b>block/positional index</b> (a trigram Bloom filter per ~1 MB block) to read only the candidate blocks — a few MB, not the whole file — so a register lookup in a huge generated header stays cheap even over a <b>network share</b>. UTF-8 files; others fall back to a whole-file line scan.</li>
<li><code>#define</code>/macro definitions are <b>not</b> captured as symbols (they'd explode the symbol index on register-map code); the names remain findable via <code>search_code</code>.</li>
<li>C/C++ references are matched by name, not compiled, so different symbols that share a name are listed together. No embeddings / semantic-meaning search. Single machine, single user.</li>
</ul>
<h2>Performance</h2>
<p class="muted">Measured on: <b>Intel Core i7-8700</b> (6 cores / 12 threads, ~2018), 32 GB RAM, Windows 11 Pro, .NET 10,
with the EcoQoS / E-core throttling opt-out in effect. A modern many-core machine will be substantially faster to build.</p>
<h3>Per-tool latency by repo (cold vs. warm)</h3>
<p>Every tool, timed two ways, over seven pinned public repos plus a huge generated C/C++ corpus (local <b>and</b> over an
SMB/UNC network share). Reproduce any row with <code>bench-matrix.ps1</code>, which fetches the public corpora on demand.</p>
<ul>
<li><b>cold</b> — a fresh CLI process per query. What you pay the <i>first</i> time a tool is used after launch: process
start + memory-map open + (for <code>find_references</code>) a cold analyzer build. The only number a one-shot CLI user sees.</li>
<li><b>warm</b> — the same query at steady state against one long-lived MCP server (how an agent uses it via Claude
Code / Codex): the index is mapped and the semantic analyzer is resident.</li>
</ul>
<p>Each query cell is <b>cold / warm</b>, median milliseconds. <code>find_references</code> latency depends on the symbol, so the
battery picks the <b>broadest</b> symbols (its worst case) and reports median and max; <b>refs 1st-call</b> is the one-time
analyzer build paid on the session's first <code>find_references</code>.</p>
<table>
<thead><tr><th>Repo</th><th>Lang</th><th>Files</th><th>Size</th><th>Build</th><th>search_code</th><th>find_definition</th><th>search_symbols</th><th>find_references (med)</th><th>find_references (max)</th><th>refs 1st-call</th></tr></thead>
<tbody>
<tr><td>requests</td><td>Python</td><td class="num">118</td><td class="num">4.9 MB</td><td class="num">1.9 s</td><td class="num">219 / 3</td><td class="num">219 / 1</td><td class="num">217 / 1</td><td class="num">855 / 4</td><td class="num">1690 / 74</td><td class="num">0.7 s</td></tr>
<tr><td>fmt</td><td>C++</td><td class="num">207</td><td class="num">3.2 MB</td><td class="num">1.6 s</td><td class="num">216 / 3</td><td class="num">217 / 2</td><td class="num">217 / 2</td><td class="num">1033 / 11</td><td class="num">1266 / 19</td><td class="num">0.7 s</td></tr>
<tr><td>EF Core <sup>1</sup></td><td>C#</td><td class="num">5,002</td><td class="num">82 MB</td><td class="num">5.3 s</td><td class="num">423 / 8</td><td class="num">218 / 2</td><td class="num">422 / 13</td><td class="num">18860 / 623</td><td class="num">44555 / 23991</td><td class="num">20.5 s</td></tr>
<tr><td>TypeScript</td><td>TS</td><td class="num">72,171</td><td class="num">349 MB</td><td class="num">26.8 s</td><td class="num">424 / 33</td><td class="num">422 / 3</td><td class="num">422 / 11</td><td class="num">1479 / 49</td><td class="num">1941 / 509</td><td class="num">1.1 s</td></tr>
<tr><td>Godot</td><td>C++</td><td class="num">9,937</td><td class="num">196 MB</td><td class="num">12.5 s</td><td class="num">434 / 11</td><td class="num">226 / 3</td><td class="num">415 / 22</td><td class="num">2491 / 43</td><td class="num">3623 / 113</td><td class="num">2.2 s</td></tr>
<tr><td>Roslyn <sup>1</sup></td><td>C#</td><td class="num">19,980</td><td class="num">372 MB</td><td class="num">16.7 s</td><td class="num">432 / 21</td><td class="num">434 / 4</td><td class="num">422 / 30</td><td class="num">46902 / 1105</td><td class="num">59792 / 18221</td><td class="num">34.3 s</td></tr>
<tr><td>LLVM</td><td>C/C++</td><td class="num">132,213</td><td class="num">1.5 GB</td><td class="num">57.9 s</td><td class="num">872 / 93</td><td class="num">639 / 7</td><td class="num">848 / 4</td><td class="num">3034 / 141</td><td class="num">4472 / 247</td><td class="num">2.0 s</td></tr>
<tr><td>generated C/C++ — <b>local</b></td><td>C/C++</td><td class="num">66,337</td><td class="num">88 GB</td><td class="num">16m 54s</td><td class="num">1018 / 11</td><td class="num">881 / 8</td><td class="num">878 / 2</td><td class="num">2329 / 459</td><td class="num">3293 / 617</td><td class="num">1.4 s</td></tr>
<tr><td>generated C/C++ — <b>SMB/UNC</b></td><td>C/C++</td><td class="num">66,337</td><td class="num">88 GB</td><td class="num">22m 11s</td><td class="num">2259 / 10</td><td class="num">2053 / 11</td><td class="num">2063 / 2</td><td class="num">11600 / 1785</td><td class="num">13486 / 3945</td><td class="num">9.6 s</td></tr>
</tbody>
</table>
<p class="muted"><sup>1</sup> The broadest names the battery picks in EF Core and Roslyn include ones declared many times over
(<code>Add</code>, <code>Contains</code>), so their <i>max</i> is the many-declarations case described below, not a typical lookup.</p>
<p>The <b>generated C/C++</b> corpus is a synthetic stress tree — <b>66,337 files / ~88 GB</b>, single source files up to
<b>1.4 GB</b> — indexed once on local disk and once over an SMB/UNC share, to show behaviour at extreme scale and across a network.</p>
<p><b>How to read it:</b></p>
<ul>
<li><code>search_code</code> / <code>find_definition</code> / <code>search_symbols</code> are index-backed: warm they answer in
<b>~1–100 ms</b> on every repo (even 132 k files, even over SMB). Cold is dominated by process start + map-open (~0.2–2 s)
— exactly what the persistent MCP server amortizes away.</li>
<li><code>find_references</code> on <b>C#</b> builds Roslyn's workspace once (the <i>refs 1st-call</i>), then keeps it resident.
Its cost then depends on the name: one with a single declaration answers in about a second warm, while a name declared many
times over (<code>Equals</code>, <code>GetEnumerator</code>) runs one whole-solution search per unrelated declaration and can
take tens of seconds — the <i>max</i> column.</li>
<li><code>find_references</code> on <b>C/C++</b> (and Python, TypeScript, …) is a name search over the index, so it costs about
what <code>search_code</code> does plus reading the matched files to skip comments and strings: well under a second warm on the
public repos. On the generated corpus, whose matched files run to gigabytes, that reading dominates — about half a second
warm on local disk, a couple of seconds over SMB.</li>
</ul>
<p>At ~25 MB/s a first index of a <b>dozens-of-GB</b> repo is on the order of minutes-to-tens-of-minutes (less on a many-core
machine), and searching it afterward uses only a few hundred MB of RAM — comfortably within a 16 GB laptop.</p>
<div class="card">
<b>Scale check (10 GB / ~1.1 million files):</b> an aggregated 10.4 GB corpus indexed in <b>7m48s</b> (22 MB/s) on the machine
above, using <b>792 MB heap / 2.2 GB peak working set</b> during a full build-plus-edit cycle. Queries stayed ~1 ms
(p95 7.9 ms). Memory stays bounded because the trigram and symbol indexes — and the change-detection ledger —
are all memory-mapped on disk, not loaded into RAM, so <b>searching</b> a repo this size uses only a few hundred MB and
<b>editing</b> it keeps just the changed files in memory. First-time build memory is tunable via
<code>CODECOMPASS_SEGMENT_MB</code> / <code>CODECOMPASS_THREADS</code>.
</div>
<h3>A closer look: <code>find_references</code> on the SMB/UNC row</h3>
<p class="muted">The same <b>88 GB / 66,337-file</b> machine-generated C/C++ corpus (332k symbols, 29 M trigram
postings, single files up to 1.4 GB, 2,925 over the 1 MB symbol cap) indexed and queried <b>over an SMB/UNC
share</b> — the SMB/UNC row above, broken out per operation (cold CLI; warm figures are in the matrix).</p>
<table>
<thead><tr><th>Operation</th><th>Over UNC (cold)</th><th style="text-align:left">Why</th></tr></thead>
<tbody>
<tr><td>First index</td><td class="num">~25 min</td><td style="text-align:left">read-bound over the wire.</td></tr>
<tr><td><code>search_code</code></td><td class="num">~2.3 s</td><td style="text-align:left">Trigram candidates + the block/positional sidecar — reads only candidate blocks, not whole files. (~10 ms warm.)</td></tr>
<tr><td><code>find_definition</code></td><td class="num">~2.1 s</td><td style="text-align:left">Memory-mapped symbol index. (~11 ms warm.)</td></tr>
<tr><td><code>search_symbols</code></td><td class="num">~2.1 s</td><td style="text-align:left">Memory-mapped symbol index. (~3 ms warm.)</td></tr>
<tr><td><code>find_references</code></td><td class="num">~12.7 s</td><td style="text-align:left">Name search: reads the matched files over the wire to skip comments and strings (~6.4 s warm); see below.</td></tr>
</tbody>
</table>
<p><b>Why <code>find_references</code> takes longer than the others:</b> for C# it builds Roslyn's model of the code
on first use (then keeps it warm); for C/C++ and other languages it reads each matched file to skip comments and strings
— over a network share, that read crosses the wire.
The text and symbol lookups only read memory-mapped indexes.</p>
<h2>Logs & troubleshooting</h2>
<p>CodeCompass keeps a self-limiting log on disk so problems can be diagnosed after the fact
without watching it live. Everything lives in one folder — run <code>codecompass logs</code> to
print the location and the current files:</p>
<pre>%LOCALAPPDATA%\CodeCompass\logs\
codecompass.log common log: startup + every warning/error, all repos
repo-<name>-<key>.log per-repo log: full detail (builds, incremental reindex, skips)</pre>
<p>Each file is <b>size-capped and rotated</b> (default 5 MB × 4 files kept), so the logs can
never grow unbounded — at most ~20 MB per file family. Warnings and errors from a repo are
mirrored up into the common log, so that one file is the first place to look when something's off.</p>
<table>
<thead><tr><th>Setting</th><th style="text-align:left">Effect</th></tr></thead>
<tbody>
<tr><td><code>CODECOMPASS_LOG_LEVEL</code></td><td style="text-align:left">off / error / warn / <b>info</b> (default) / debug</td></tr>
<tr><td><code>CODECOMPASS_LOG</code>=0</td><td style="text-align:left">disable logging entirely</td></tr>
<tr><td><code>CODECOMPASS_LOG_DIR</code></td><td style="text-align:left">write logs somewhere other than the default folder</td></tr>
<tr><td><code>CODECOMPASS_LOG_MAX_MB</code></td><td style="text-align:left">rotate each file at this size (default 5)</td></tr>
<tr><td><code>CODECOMPASS_LOG_KEEP</code></td><td style="text-align:left">how many rotated files to keep (default 3)</td></tr>
</tbody>
</table>
<p class="muted">Logging never throws and never blocks indexing or search — if the log can't be
written it's silently skipped. Set <code>CODECOMPASS_LOG_LEVEL=debug</code> to record per-file skips
and detailed build stats when chasing a specific issue.</p>
<footer>
CodeCompass — local code intelligence for AI agents. Fully local, no GPU, minimal token footprint (5 MCP tools).
Perf figures measured on the machine noted above; your numbers will vary with hardware and codebase.
</footer>
</div></body></html>