Repository navigation
release: 0.41.0 - #128
Merged
Merged
release: 0.41.0#128
Conversation
Codex ships a plugin system built on the Agent Plugins open standard: a plugin.json manifest and an mcp.json beside it, both validated against agent-plugins.org/schemas/1.0.0. Claude Code wants the same content one directory down, in .claude-plugin/, with the MCP config dot-prefixed. The formats differ only in where the manifest sits and how presentation fields nest, so the bundle carries a manifest pair per client and a single copy of the skill. A second copy of SKILL.md is the drift this project exists to prevent, and the test pins it against canonical either way. Both clients can now install straight from the repository: the root carries a marketplace catalog for each, and both resolve the plugin at ./plugins/agentsmesh-lessons, which is how openai/plugins lays its own catalog out. Two things deliberately left out. The legacy .codex-plugin/plugin.json that OpenAI's own plugins still ship is a documented compatibility fallback, and carrying both manifests would leave Codex picking between two sources of truth that cannot be tested here. Hooks stay out for the reason they always did: Codex has no failure-only lifecycle event, and every hook call pays a fresh npx resolution the MCP server pays once. The manifests carry versions, so sync-server-json.ts becomes sync-release-versions.ts and bumps all three during `changeset version`.
`codex plugin marketplace add` refused the whole catalog: unknown variant `NONE`, expected `ON_INSTALL` or `ON_USE` The field is optional and this plugin authenticates nothing, so it has to be absent rather than set to a value that spells "none". Found by running the real CLI against the catalog, which no schema check would have caught.
`init --lessons` writes a blocking recall/capture rule into the user's instruction file, so it reaches every tool as a root rule. A plugin has no such reach: it may ship skills, hooks and servers, never that file. The bundle also ships no hooks, because each invocation would pay a fresh npx resolution the server pays once per session. So a plugin-only install had no standing instruction at all. The only always-on signal was the skill's one-line description, which is an advisory "use when" pointer — recall became discretionary and the completion receipt was never demanded. Worst hit was start-of-task recall, the call that surfaces universal rules: nothing pointed the model at it, because the skill's own trigger is "about to edit a file". `instructions` is the one channel a server has for standing text, and it costs nothing per call. Same three obligations as LESSONS_PROCEDURAL_RULE, in tool vocabulary rather than shell, since a client reaching lessons this way may have no shell. Asserted on the wire, not just on the constant: the value is useless if the SDK does not serialize it into the result. The guide claimed the plugin was a lighter route to the same thing. It now says an instruction is followed by choice where a hook simply fires.
From the repo-wide over-engineering audit, the findings with the highest
value and the lowest risk. Every one verified against the tree before
touching it, not taken on the scanner's word.
Dead, with no caller outside its own tests:
- hasInterpolation / usesCursorSensitiveInterpolation, exported and
tested since the first commit, never imported by any target. No doc
comment claimed a reason to keep them.
- narrowDiscoveredForImplicitPick, fully superseded by
narrowDiscoveredForInstallScope, which has two production callers.
Its dedicated test file went with it.
- copilot/hook-entry.ts, a five-line pass-through that re-exported a
core function under its original name. Its two callers now import
the core function directly.
Duplicated:
- Fifteen targets each spelled out the same four-line ignore generator,
differing only in a path constant. One ignoreOutput(path) factory in
the catalog, in the shape NO_OUTPUTS already set. The four targets
that gate the file by scope keep their own generator.
- The pack writer and the pack merger each repeated the same
copy-into-subdirectory loop for rules, commands and agents. Six
copies, one helper. Skills stay separate: they nest a directory per
skill and carry supporting files.
Also dropped four knip entry patterns pointing at files that no longer
exist, which silenced every stale-entry hint it was printing.
No behaviour change. generate --check is clean, so every generated
artifact is byte-identical.
Merging master back brings package.json to the released version, but the plugin manifests live only on develop, so they kept 0.39.0 and the bundle test failed — the guard doing its job. `scripts/sync-release-versions.ts` is the fix, and from the next release the version PR runs it itself.
Caught by the pre-release regression sweep, in the change I shipped two commits ago. The instructions field went out unconditionally, but this server is not only a lessons server: it carries the config tools and the README advertises it on its own. Most people who wire it up never ran `init --lessons`. Those users were getting a BLOCKING contract that made three false promises at once. It called `.agentsmesh/lessons/lessons.json` canonical when the file was absent, pointed at a `lessons` skill they never installed, and required a query before every edit that could only ever return no matches. Obeying the capture half was worse: a valid lessons_add scaffolds a graph and a capture log into a repository that never asked for either, untracked. The text is state-aware now. Where lessons exist — by graph or by config, so a capture-only bootstrap counts — the contract is unchanged. Where they do not, the server describes what it offers, names the fields a first capture needs, and mandates nothing. Verified on the wire in both states: 441 characters and no mandate in a project without lessons, the 749-character contract in one with them.
…ostile graphs Fixes every gap from the Staff audit of the lessons system. Teams - Field-level three-way merge driver; generate and init set up the per-clone driver (refused when its program is not on PATH, custom drivers kept); new "lessons resolve" repairs a conflicted graph. - check and generate --check fail on an unreadable graph (conflict, corrupt, newer version); a normal generate warns. - Process lock uses owner tokens: no eviction or release of another holder's lock; lessons lock stale after 60 s. - A file_glob is dead only with git history proof; fresh clones no longer prune pending lessons. - Hooks run "npx --no --offline agentsmesh" when the project depends on it, plus a team hint when it does not. Hosts - Recall hook answers Gemini BeforeAgent, Cursor and Copilot sessionStart/postToolUseFailure (Copilot exit code 2 passed through); descriptors (and plugins) declare hookContextEvents. - Codex apply_patch recall, per-subagent dedup, lessons root found from payload cwd / CLAUDE_PROJECT_DIR / subdirectory (hook and MCP). Capture and effectiveness - Failure text from Claude Code's error field; class past "Exit code N"; keyed on the real program; interrupts not recorded. - Outcome log on by default (outcomeLog / AGENTSMESH_LESSONS_OUTCOME_LOG); a miss needs a same-session failure within 30 min on the lesson's own trigger. Effectiveness is memoized and scoped to ranked lessons: a full 5000-event log adds ~20 ms per hook call, not ~130 ms. - --trigger-file outside the project is rejected. Safety - Recalled rules fenced in <recalled-lessons>, one line each, with ids; CLI and MCP answers capped at 32,000 chars; only printed rules are marked seen; --always uses the always-on budget. - file_glob uses a linear matcher over a safe subset on every path (recall, validate, prune, capture, effectiveness); UNSAFE_GLOB_PATTERN. - recallLimit/recallMaxTokens ceilings; legacy migration stays inside .agentsmesh/lessons. Docs, changeset, and regenerated skill copies included.
Captures the reusable rules found while fixing the lessons audit: safe glob matching on every path, git-proof dead globs, merge-driver and merge-field rules, host hook shapes, hook latency, printed-only dedup, legacy path containment, lock ownership, temp git repo isolation, orphaned vitest workers, stash-free red proofs, and zsh echo of JSON. Supersedes the two lessons that described behavior that never existed.
Applies the verified ponytail review of the unpushed work. Behavior is unchanged except one wording: recall on an unreadable graph now prints the shared graph-problem message (CLI and MCP), so a merge conflict always points at "lessons resolve". - lock: existsSync, one evictOwners, backoff folded into process-lock - globs: drop test-only helpers and a repeated length check; inline deadFileGlobIds; shared liveness and timing test helpers - hook: isOlderVersion replaces a full semver parser; drop test-only emitRecall, optionalPreface and the env parameter; one findUp - merge/effectiveness: one git runner, mergeScalar reuse, recurringFailure in the failure hook, isIneffective in one place, configFlag reuse, AddLessonOptions back to its published shape, 17 file-local exports - cli/mcp/install: shared pack copiers, one degraded-recall warning, isCaptureRejection, lessonsShow caps itself, inlined one-caller helpers - targets: recall projection only in the engine; reverse maps and lint lists derived from the forward maps; shared canonical test factory Skipped on purpose: stripVTControlCharacters (strips more), v1 absentGraph (emptyGraph would upgrade merges to v2), flat memo Map (2.5x slower), the scripts core/wrapper split (repo convention), the rich-plugin fixture test. Verified: full suite + coverage floor, e2e, lint, typecheck, knip, and an old-vs-new dist run of one scenario with identical transcripts and trees except the intended wording and volatile pack metadata.
Five exploratory QA sessions on the built CLI found 4 high and 17 medium defects in the unpushed lessons work; all are fixed test-first. High - Hook never breaks the host: best-effort log writes, readers skip bad lines, guarded lock read, catch-all in doHook (it used to exit 1 and lose the lesson for the session). - A merge driver git could not start silently dropped the other branch's lessons: validate/check/generate --check now flag an unmerged lessons.json missing the union; setup ignores npx-cache and node_modules/.bin PATH entries and checks the repo root for npx. - MCP lessons tools never create a graph in home or outside a project (lessons folder, then agentsmesh.yaml, then git repo root); the hook is silent and CLI add/import-md refuse in the home folder. - CLI refuses extra positionals (an unquoted rule kept only its first word). Medium - Single-value flags given twice, read-only tool failures, one merged recurrence warning per patch within the cap, prompt-less graph notices, fence against invisible and look-alike characters, graph-problem messages everywhere (SCHEMA_INVALID), resolve after a hand fix, future locks stale, MCP always/no_dedup/session, lessons_show by id, typed write refusals as VALIDATION_FAILED, negated globs broad, upsert changes reported, subfolder root resolution, dead command patterns per the documented contract, pack hooks never dropped (extends still override per event), a writer that lost the lock refuses to save. Verified: full suite + coverage floor, e2e, lint, typecheck, knip, website build, generate --check, and a replay of every confirmed defect.
…nt drops - Re-installing the same source updates its pack instead of failing or adding a second one: the pack named by --name, else the one with the same features, else the single whole-source pack (same source, target and as). A whole-source re-install replaces the pack like refresh; picks merge. The pack keeps its name, and --dry-run names the pack it would update. - A --name held by another source's pack fails with a message about --name. - Any skills/<dir>/SKILL.md marks a skill pack, so skill folders that are not kebab-case install together with rules, README and LICENSE. - Warn, naming each file, when mcp.json, hooks.yaml, permissions.yaml or ignore sit at the root of a source without .agentsmesh/.
generate rewrote generated_at in .agentsmesh/.lock on every run, even when it printed "Nothing changed", so the git tree was dirty after each run and every teammate's generate showed a lock diff. writeLockFile now compares the new checksums, extends, packs and outputs with the lock on disk and skips the write when they match. generated_at, generated_by and lib_version are not compared: they differ per run, user and version, so comparing them would bring the churn back. They now describe the last run that changed the lock. check, generate --check and watch are unchanged.
- A value flag given no value is exit 2 ("--flag needs a value"); values
that start with -- go as --flag=value; `lessons help [sub]` works.
- prune --cap uses the strict positive-integer check query already had.
- add: bad --scope, a topic id that is not kebab-case, a new topic without
a non-blank summary, and an unsafe --trigger-file glob are exit 2 with
messages that name the flag; write refusals are exit 2 with a plain
message; internal function names are gone from errors; evidence refs
are stored once; the rule limit and recall clamp count characters.
- --trigger-file is stored in one normalized, project-relative form;
../ paths, existing folders and the project root (also through a
symlink) are refused, on the CLI and in MCP lessons_add.
- \u{...} in a command pattern is rejected, as documented.
- query names a failed legacy migration; import-md --migrated-at must be a
real date; env switches accept true/false/yes/no/on/off; config warns on
non-boolean switches and non-object files.
- show prints the rationale, journal marks retired lessons, validate
--json names the error codes, and the docs/help text are corrected.
tsconfig.json leaves tests out and vitest strips types, so fixtures drifted from the real types unseen: - lessons renderer tests: stats fixtures gain effectiveness/hasOutcomeLog (the effectiveness test also sets misses/failingActions and asserts them), the show fixture uses `subject`, the help result has data: null. - lessons command tests: assert the seeded lesson exists (noUncheckedIndexedAccess). - lessonsInit helper: drop the `as` cast that hid the missing rootRuleMerged, mergeDriver and recallHookTeamHint fields; defaults print nothing. - e2e reference helpers: TargetName is now the catalog's BuiltinTargetId (codebuff, openhands and kimi-code were missing), a .ts import became .js, and the frontmatter capture is checked before use. tsconfig.tests.json type-checks src plus every test file, except tests/consumer-smoke (own tsconfig) and a baseline of 274 files that still have errors. New and fixed files are always checked; `typecheck:tests` runs it and passes today.
Hook: drain stdin past the 1 MB cap (no broken pipe), read JSON with a BOM, do nothing for an unknown event name (a missing one stays a tool call), give Copilot's VS Code SessionStart task recall, take the project from the touched file first, match NFD paths against NFC globs, name what the prompt caps hid, count recurrences over the last 24 hours, skip recording Cursor permission_denied, drop the file-glob hint from command nudges, and re-read the session dedup store under a short lock. Logs and files: start a new line after a cut line, cap logs by bytes and read only a bounded tail, refuse to write a read-only lessons.json and keep its mode, name a lock path that is a file, ignore a BOM in lessons.json and config.json, say once who holds a busy lock, and sweep old temp files and stale lock folders. Merge: keep a trigger or topic one branch deleted, refuse to save a marker-rebuilt result with errors and warn without a diff3 base, count after renames, name the next git step per operation, fall back to the bare driver without npx, compare the wired recall command with the project, and stop repeating the custom-driver line on every generate.
run-install-execute.ts had grown to 252 lines, over the 200-line limit. The selection half (elevated-artifact consent gate, resource pools, conflict resolution, features, pick and entry name) moves to run-install-selection.ts as selectInstall(); run-install-execute.ts keeps RunInstallExecuteArgs, InstallExecuteResult and the write step (extends vs pack, dry-run, post-install generate). No behavior change: the old and new builds give identical output and files across 10 install scenarios.
Three pages said `agentsmesh merge` drops the outputs map, so `check` skips output verification until the next generate. merge keeps the map: it combines both branches' entries, and your branch's hash wins where both recorded a file. Checked end to end: after a real .lock conflict, `check` reports the file only the other branch changed as modified (exit 1) until `agentsmesh generate`; only a merge of two locks without an outputs map leaves the key out and gets the "skipped" notice.
check told every drift to "run 'agentsmesh merge' ... or 'agentsmesh generate --force'". After a lock conflict resolved with merge, that sent users to the wrong command for plain generated-output drift. - output drift: run 'agentsmesh generate' (replaces hand edits) - canonical drift: run 'agentsmesh generate' (merge would hide it) - locked features: revert, or 'agentsmesh generate --force' - a lock with git conflict markers is now reported as a conflict (JSON lockConflict: true) and points to 'agentsmesh merge', instead of "Not initialized for collaboration"
The MCP check tool now returns lockConflict: true when .agentsmesh/.lock has git conflict markers, like 'agentsmesh check --json', so an agent runs 'agentsmesh merge' instead of treating the project as never generated. The marker check moves to core/check/lock-conflict.ts and is shared by the CLI and MCP.
The programmatic check() now reports lockConflict: true when .agentsmesh/.lock has git conflict markers, so API users can tell a conflicted lock (fix: agentsmesh merge) from a missing one. The engine report is now the one place that knows this: the CLI and the MCP check tool read report.lockConflict. The report and options interfaces move to lock-sync-types.ts (re-exported) to keep lock-sync.ts under 200 lines.
… helper - runCheck built the same empty result checkLockSync already returns without a lock; it now maps the report in one place. - lock-conflict.ts had one caller left; lock-sync.ts reads the lock with readFileSafe and checks the markers inline. - The CLI and MCP conflict tests keep only the conflict case; the engine test covers the other cases. check output (text, --json, exit codes) is byte-identical for no lock, a conflicted lock, a conflicted lock with a BOM, and a readable lock.
- check now delegates to the CLI runCheck, like generate delegates to runGenerate. An unreadable .agentsmesh/lessons/lessons.json fails 'agentsmesh check'; the MCP result now reports it in a new lessonsGraphError field (the CLI JSON error text, null otherwise) instead of looking clean. - generate sets lockfileUpdated from the real lock write. writeLockFile returns whether it wrote, and runGenerate passes it up as lockWritten, so a no-op run that leaves the lock alone reports false. - The MCP check unit tests move out of orchestrate.test.ts into orchestrate-check.test.ts; the real-project MCP check test is renamed orchestrate-check-project.test.ts.
- The lessons resolve and graph-problem tests ran `git merge` without a git identity. A CI runner has none, so git stopped before merging and the old `status !== 0` check passed the wrong failure. They now use a shared tryGit helper with the fixed test identity and assert a real content conflict. - launcher tests join PATH with path.delimiter on the host platform (a Windows path has a drive-letter colon). - merge-driver-setup-launcher compares with realpathSync.native, which expands Windows 8.3 short names the way git does. - seen-cache-concurrency runs its workers with `node --import tsx` and imports through a file URL (Windows has only tsx.cmd, and a C:\ path is not an ESM specifier).
On Windows, rmdir of an owner marker that another process is removing at the same moment fails with EPERM or EBUSY for a short time. removeOwner only accepted ENOENT, so acquireProcessLock crashed a waiting command (seen in the cross-process stress test on a Windows runner). It now retries these codes with a short backoff, like renameWithRetry, returns false once the marker is gone, and still throws an error that does not clear after 5 attempts.
…tore lock codecov/patch on the release PR was at 99.01% of changed lines against a 99.29% target. These tests cover the missed lines that are real paths: - tryAcquire throws a mkdir error other than EEXIST; - evicting an older-version lock throws a rename error other than ENOENT; - processIdentity reads the start time with ps (the macOS probe; Linux runners have ps too); - a seen-store commit writes anyway after a lock that never frees has waited its 2 seconds, and leaves that lock alone.
…seen store
The second CI round found more of the same Windows race: a folder or file
another process is removing or reading fails mkdir, rmdir or rename with
EPERM, EACCES or EBUSY for a moment ("delete pending").
- New transient-fs.ts: retryTransient and retryTransientSync retry those
codes with a short bounded backoff and still throw after 5 attempts.
- tryAcquire retries the lock-folder mkdir (it crashed the stress test);
removeOwner uses the shared helper instead of its own loop.
- seen-store: takeLock treats a transient error as a busy lock instead of
writing unlocked, and writeSeenStore retries its rename instead of
silently dropping the write (a 12-process test lost an id on Windows).
…lock The third CI round hit the same Windows race in one more call: readdir of a lock folder another process was removing failed with EPERM while inspecting it. Rather than wrap each call, acquireProcessLock now treats a transient error (EPERM, EACCES, EBUSY) from claiming or inspecting the shared lock path like a busy lock: it waits briefly and tries again, and 5 in a row still throw. The claim's own mkdir retry goes, so there is one layer. Eviction stays outside that catch, so a real eviction error still surfaces at once (process-lock-recovery test). Only removeOwner keeps its bounded rmdir retry, because it races other evictors on the shared owner marker.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cuts 0.41.0. Twelve changesets are pending: four minor, eight patch.
What ships
Lessons for teams, more tools and hostile graphs (
steady-lessons-hold,calm-lessons-guard). Two branches that both capture lessons now merge cleanly through a field-by-field git merge driver, withagentsmesh lessons resolvefor leftover conflict markers. The recall hook now covers Gemini CLI, Cursor, GitHub Copilot and Codexapply_patch. File triggers use a linear-time safe glob matcher.checkandgenerate --checkfail whenlessons.jsoncannot be read.MCP clients get the lessons contract when they connect (
plain-moons-greet). Where a project has no lessons, the server says so and mandates nothing.Upgrade notes (
upgrade-notes). Seven things to check after upgrading: rungenerateonce, runlessons validatefor globs that no longer match, Gemini hook matchers, new Copilot and Cursor hook events, the outcome log default, empty flag values, and packs replaced on re-install.Fixes (patch):
checkgives the fix that matches the drift. A conflicted.agentsmesh/.lockis reported aslockConflictin the CLI, the MCP tool and the programmaticcheck().generateleaves.agentsmesh/.lockbyte-identical when nothing changed.checktool reports an unreadable lessons graph (lessonsGraphError), and MCPgeneratereports the real lock write.agentsmesh lessonsrefuses bad input clearly; hook, merge driver and log edge cases are fixed.EPERM, and parallel lessons recalls no longer lose a session dedup entry.Also in this range, outside the npm package: one plugin bundle that installs in both Codex and Claude Code.
Verification
generate --check,check,lessons validatenpm pack --dry-runpnpm audit --prodNotes
The first CI run on this range failed on tests that had only run on macOS before. Fixed in
bb1d46bb(merge tests rangit mergewithout a git identity, which CI runners lack; Windows PATH, 8.3 path and tsx spawn portability in four tests) andcab70b64(the Windows lock race above).117795edadds tests for the error pathscodecov/patchflagged. The second run then showed the same Windows race in two more places (claiming a lock folder, and the lessons seen store losing an id);2abb9a92moves the retry into one shared helper,transient-fs.ts. The third run hit it once more inreaddir, so2c18bea8handles it once inacquireProcessLock: a short Windows error while claiming or inspecting the lock is treated like a busy lock. Eviction errors still surface at once.Master was merged back first in an earlier step, so the four changesets 0.40.0 already consumed are not counted twice.
enginessays Node>=20, but the CI matrix tests Node 22 and 24 only. This was already true in 0.40.0; worth deciding separately (add Node 20 to CI or raiseengines, since Node 20 is end of life).