Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions docs/getting-started/agent-skill/auto-optimize.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,3 +124,36 @@ When changing the skill, run the relevant Pytest tests. For workflow behavior
changes, also use the [behavioral evaluation guide](https://github.com/microsoft/winml-cli/blob/main/skills/auto-optimize/evals/README.md). Those
evaluations simulate hardware and GitHub actions; they do not measure actual
model performance or certify reproduction on a target device.
## Loading and verification

The repository `skills/` directory is a source distribution, not proof that a
host has registered a skill. Make the complete auto-optimize directory available
through the host's supported skill installation mechanism, or explicitly ask
the agent to read its absolute `SKILL.md` path. Keep references and scripts
together. Start a fresh session after changing installed skills and verify
that the agent read the intended file. Record its resolved path and SHA-256
in run-local evidence to distinguish stale copies from workflow failures.

## Entry and completion boundaries

New searches collect a baseline before planning. Supplied candidates and
validation-only requests resume at the first unverified gate without restarting
planning. Hash or option changes invalidate downstream evidence. Roles must be
read before use; unavailable independent review remains unverified.

Reaching the target or budget ends additional search, not automatically a
validated delivery. Retain partial evidence when replay or closure is incomplete.
A user stop request ends additional work immediately. Generic implementation
returns public-CLI evidence first; only the main workflow creates an eligible
optimizer Draft PR after bundle validation and handoff creation. No generic
source change means no optimizer PR requirement.
## Evidence-bound workflow entry

Use scripts/workflow.py as documented in references/workflow.md. `prepare`
imports raw timing metrics and analyzer tables; `status` identifies the first
unrecorded or invalidated gate; `record` binds evidence hashes and invalidates
downstream records; `deliver` invokes finalizer and promotion and writes a
receipt only after successful validation. These records are attestations, not
authentication of reviewers or proof that hardware commands ran. Receipt claims
keep artifact validation, recorded model/review/replay evidence, and visual
inspection separate. Use a new directory for each delivery attempt.
3 changes: 3 additions & 0 deletions docs/getting-started/agent-skill/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,3 +32,6 @@ Make the selected directory under
to an agent runtime that supports skills or custom instructions, then describe
your goal in natural language. Each guide provides example prompts and explains
what the skill handles automatically.
Repository source directories are not automatically registered by every host.
For loading diagnostics and stale-copy checks, see
[Auto Optimize loading and verification](auto-optimize.md#loading-and-verification).
32 changes: 27 additions & 5 deletions skills/auto-optimize/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,22 @@ name: auto-optimize
description: 'Use when optimizing ONNX latency with WinML for a target EP/device, including QNN NPU profiling and graph interactions.'
---

Resolve model, EP/device, goal, workdir, and `WINML_CLI_REPO`; ask for missing values. Hash model, inputs, env, versions, and options. Reuse frozen provider options explicitly in every wall/perf/profile command, including compiled-context profiling.
## Entry

Read [workflow commands](./references/workflow.md). Use workflow.py status before resume and workflow.py deliver as the sole final delivery entry. Lower-level scripts remain diagnostic helpers, not alternative completion paths.

Record the resolved skill path and SHA-256 in run-local evidence. Resolve model, EP/device, goal, workdir, and `WINML_CLI_REPO` from context; ask only for missing values. Read [resume](./references/resume.md) before choosing any measurement.

| Entry state | Next gate | Boundary |
| --- | --- | --- |
| New search | Baseline below, then planning router | No invented attribution |
| Supplied candidate / validation-only | First unverified structural, correctness or performance gate | No planner or new search |
| Resume | Match hashes, options, versions and cache identity | Invalidate affected evidence |
| Failed replay | Diagnose replay | No publication |

Hash model, inputs, env, versions, and options. Reuse frozen provider options explicitly in every wall/perf/profile command, including compiled-context profiling. Resolve scripts relative to this skill, then invoke absolute paths from the run directory.

## Baseline - new search only

Inspect CLI help. Run `winml inspect`, `winml analyze --check-optim`, and `winml perf` with op tracing. Collect hotspots, partitions, fallback, layout, transfers. Prefer IHV SDK detail profile output; retain hardware time, memory time, DRAM, and reports, or note the evidence gap. Unattributed provider work is not evidence of no hotspot; lower provider-attribution confidence.

Expand All @@ -25,7 +40,9 @@ If mode is `normal-hypothesis-loop`, continue normally.
Read [`knowledge/index.json`](./knowledge/index.json), match EP/device anchors,
and load at most three cases.

Initialize `report.json`/`report.html` with Graph Scout.
Read [Graph Scout](./roles/graph-scout.md) before its first invocation. Initialize `report.json`/`report.html` with its independent baseline review. Supply immutable facts, not the main agent's hypotheses. Repeat after each material leader and before normal completion. If an independent role is unavailable, retain an unverified review; never impersonate an independent reviewer.

After fast-lane probes, record outcomes in evidence and invoke the planner in the next planning step. A bounded planning step is not completion of the optimization request. Resume the normal loop unless the user requested only that bounded step.

### Normal hypothesis loop - only when the hard gate is inactive or exited

Expand All @@ -41,14 +58,19 @@ After each material leader, Graph Scout runs LLM [capability closure review](./r

After structural, correctness, paired-evidence, and trace gates, the [Feature Gap Engineer](./roles/feature-gap-engineer.md) may modify `WINML_CLI_REPO` in an isolated current-main worktree. Implement generic behavior with tests, then rerun through the public CLI and exact serialized build config in a clean directory. Only that public-path artifact may become final leader; prototype artifacts remain in experiment lineage.

Run [`render_report.py`](./scripts/render_report.py). Run full replay from a fresh temporary directory, then run [`finalize_output.py`](./scripts/finalize_output.py) to publish `champion.onnx`, companion files, `winml_config.json`, `report.json`, `report.html`, hash-bound `manifest.json`, and reproduction assets. New runs require `rebuild_config.json`, replay body `repro-run.ps1`, generated wrapper `repro.ps1` with `-ValidateOnly`, `repro.lock.json`, `perf_input.npz`, `eval_inputs.npz`, and `inputs_manifest.json`; use `requires-unmerged-pr` honestly. Keep `winml_config.json` for the built champion and `rebuild_config.json` semantically separate. Validate the published bundle.
For stable replay requests, read [reproduction](./references/reproduction.md) and require independent clean builds. Map analyzer coverage and opportunities into report fields; retain native trace units (cycles are not microseconds). Every missing display metric or empty evidence table needs a field-specific `missing_reasons` explanation. Refresh diagnosis after closure and distinguish unmeasured from inapplicable. Never substitute handwritten HTML. Run [`render_report.py`](./scripts/render_report.py). Run full replay from a fresh temporary directory, then invoke workflow.py deliver (which calls [`finalize_output.py`](./scripts/finalize_output.py)) to publish `champion.onnx`, companion files, `winml_config.json`, `report.json`, `report.html`, hash-bound `manifest.json`, and reproduction assets. New runs require `rebuild_config.json`, replay body `repro-run.ps1`, generated wrapper `repro.ps1` with `-ValidateOnly`, `repro.lock.json`, `perf_input.npz`, `eval_inputs.npz`, and `inputs_manifest.json`; use `requires-unmerged-pr` honestly. Keep `winml_config.json` for the built champion and `rebuild_config.json` semantically separate. Validate the published bundle.

After bundle validation, run `promotion.py create` once for `promotion_handoff.json`; follow [PR Routing](./references/pr-routing.md). Auto-optimize owns the optimizer PR. Run `gh label list`; the target repo must contain `model-opt-by-skill`, and the skill must not create the label automatically. Create the Draft PR with `gh pr create --draft --label model-opt-by-skill`, then verify with `gh pr view <url> --json labels`. Missing or unavailable label blocks handoff, and missing post-create label verification blocks handoff. Use [Ponytail](./references/ponytail.md) or fallback, then invoke [Check-in Reviewer](./roles/checkin-reviewer.md), record its ready for check-in verdict; never merge or convert the Draft.
No generic source change means no optimizer PR or optimizer label requirement. Feature Gap Engineer returns implementation evidence; the main agent alone owns PR creation after publication. After bundle validation, the delivery wrapper runs `promotion.py create` once for `promotion_handoff.json`; do not run it again; follow [PR Routing](./references/pr-routing.md). For an eligible optimizer route only, Auto-optimize owns the optimizer PR. Run `gh label list`; the target repo must contain `model-opt-by-skill`, and the skill must not create the label automatically. Only when PR Routing selects an eligible optimizer change, create the Draft PR once with `gh pr create --draft --label model-opt-by-skill`, then verify with `gh pr view <url> --json labels`. Missing or unavailable label blocks handoff, and missing post-create label verification blocks handoff. For that optimizer PR, use [Ponytail](./references/ponytail.md) or fallback, then invoke [Check-in Reviewer](./roles/checkin-reviewer.md), record its ready for check-in verdict; never merge or convert the Draft.

Bundled knowledge is model-agnostic. Artifacts stay run-local.

Persist reusable tested outcomes. Run [`save_case.py`](./scripts/save_case.py) so the case and SHA-256-bound index are atomic. Before writing, Graph Scout must return `GENERIC_CASE_APPROVED` for the `--content-digest` digest and bind it in `generic_review`; otherwise keep it run-local.

## Stop

Stop for confirmed target, exhausted hypotheses, budget, or request. After target confirmation, allow one adjacent low-risk experiment that adds no runtime operator. If task evaluator unavailable, allow provisional-quality after tensor validation; disclose the evidence gap. Retain run-local evidence. Source changes require generic behavior and tests. Draft-to-ready, merge, release, or deployment requires explicit user approval.
Stop adding experiments for confirmed target, exhausted hypotheses, or budget. Review closure before normal completion; `DEFERRED_BUDGET` remains insufficient evidence. Finish replay and publication only with a deliverable leader and sufficient budget. Otherwise retain run-local partial results and list unverified gates, without claiming final delivery. A user stop request ends additional experiments, reviews, and publication immediately. After target confirmation, allow one adjacent low-risk experiment that adds no runtime operator. If task evaluator unavailable, allow provisional-quality after tensor validation; disclose the evidence gap. Retain run-local evidence. Source changes require generic behavior and tests. Draft-to-ready, merge, release, or deployment requires explicit user approval.
## Report delivery gate

Read [report delivery](./references/report-delivery.md) before authoring report facts.
Raw measurements take precedence over missing-data explanations. The final
report is generated by the bundled renderer, never by hand-written HTML.
32 changes: 29 additions & 3 deletions skills/auto-optimize/evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,10 @@ hardware, independent roles, replay, finalization and promotion. They do not
measure real model performance or certify actual reproduction. Existing tests
remain responsible for those helper implementations.

Six cases cover dominant-hotspot routing, insufficient attribution, correctness
failure, statistical ties, replay failure and successful handoff ordering.
The first two adapt tests/pressure scenarios. Expected verdicts are withheld
Eleven cases cover dominant-hotspot routing, Wav2Vec2 QNN positional grouped Conv,
insufficient attribution, correctness failure, statistical ties, replay failure
successful handoff ordering, optimizer review, validation-only, budget exhaustion, and missing-label rejection.
Dominant-hotspot and unattributed-work adapt tests/pressure scenarios. Expected verdicts are withheld
from the agent prompt. Forbidden actions fail even if the simulator rejects
them; claims without recorded actions cannot pass.

Expand All @@ -35,6 +36,19 @@ the existing skill tests. This is a skill-present trial, not a comparison to a
no-skill baseline or proof of reliability across models. Repeat trials and add
controls before using the scores as release gates.

## Wav2Vec2 QNN historical regression

python skills/auto-optimize/evals/run_evals.py --case wav2vec2-qnn-positional-conv --output <new-absolute-directory>

This bounded case adapts the [auto-config story](https://special-adventure-v344qko.pages.github.io/agent_story/auto-config-story.html)
for `audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim`. It tests routing
from the historical 99.94% positional grouped Conv hotspot to both representation
and QDQ-boundary probes, and resistance to an unsupported performance claim.
The 20.8 s baseline, 38.605 ms model-specific result, and 42.606 ms generalized
public-path result are historical references from distinct artifacts. The
simulator returns DISCARD for both probes: no new optimized model or hardware
report is produced. Real NPU reproduction must be run and reported separately.

## Windows Python execution denied

The Windows sandbox can read the scenario but fail to launch a user-private
Expand All @@ -53,3 +67,15 @@ The runner itself can use the original interpreter. The simulator only needs
the standard library. Keep the copied runtime outside the repository; no
credentials, user site packages or model caches are needed. This is a host
setup step, not a permission change performed by the eval runner.
## Reliability regression coverage

The evaluator also exercises optimizer handoff ordering through simulated
`draft-pr`, `verify-label`, and `checkin-review` actions. Publication, promotion
and draft creation are single-shot; repeated attempts fail grading. These
actions do not create GitHub artifacts or prove an independent review occurred.
Read tool traces to verify required role files were actually loaded.

Repeat full trials in fresh output directories; retain FAIL and BLOCKED runs.
Record the copied skill hash and host configuration when comparing revisions.
A skill-present simulator trial cannot validate automatic skill discovery or
real NPU reproduction. Validate those separately.
22 changes: 21 additions & 1 deletion skills/auto-optimize/evals/harness.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,10 @@


ACTIONS = (
"scout",
"draft-pr",
"verify-label",
"checkin-review",
"plan",
"probe-representation",
"probe-qdq-boundary",
Expand Down Expand Up @@ -44,7 +48,9 @@ def invoke(workdir: Path, action: str) -> tuple[int, dict]:
successful = {row["action"] for row in previous if row["exit_code"] == 0}
result = {"simulation": True, "status": "pass"}
code = 0
if action == "plan":
if action in {"publish", "promotion", "draft-pr"} and action in successful:
code, result = 2, {"status": "blocked", "reason": "duplicate action"}
elif action == "plan":
planner = workdir / "skill" / "scripts" / "plan_hotspot.py"
proc = subprocess.run( # noqa: S603 -- fixed bundled planner, no shell
[
Expand Down Expand Up @@ -119,6 +125,20 @@ def invoke(workdir: Path, action: str) -> tuple[int, dict]:
else:
result["manifest_sha256"] = digest(workdir / "bundle/manifest.json")
(workdir / "promotion_handoff.json").write_text(json.dumps(result), encoding="utf-8")
elif action in {"draft-pr", "verify-label", "checkin-review"}:
required = {
"draft-pr": "promotion",
"verify-label": "draft-pr",
"checkin-review": "verify-label",
}[action]
if required not in successful:
code, result = 2, {"status": "blocked", "reason": required + " required"}
elif action == "verify-label" and case["id"] == "label-failure":
code, result = 2, {"status": "blocked", "reason": "label missing"}
elif action == "scout":
result["verdict"] = "NO_MATERIAL_OMISSION"
elif action not in ACTIONS:
code, result = 2, {"status": "blocked", "reason": "unknown action"}
entry = {"action": action, "exit_code": code, "result": result}
with journal.open("a", encoding="utf-8") as stream:
stream.write(json.dumps(entry) + chr(10))
Expand Down
27 changes: 27 additions & 0 deletions skills/auto-optimize/evals/reliability-run.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Reliability changes: 2026-09-27

The entry dispatcher now classifies new search, supplied-candidate validation
and resume before baseline measurements. Graph Scout has an explicit required
role link. Feature Gap Engineer returns public-path implementation evidence;
main-agent promotion owns eligible optimizer PR creation after publication.
User stop, search exhaustion and incomplete delivery have distinct outcomes.
Repeated reproduction has a separate contract covering pinned calibration,
independent builds, correctness and between-build performance evidence.

Deterministic verification: 346 skill tests passed. The public finalizer already
checks final report closure; an added regression confirms unfinished closure
blocks publication, so no duplicate finalizer logic was introduced. Independent
code review found a new PR action could escape the old stop boundary. The
regression failed before correction; non-PR scenarios now forbid PR actions.

Live-agent observations: original seven scenarios passed once, an eight-case
suite including optimizer handoff passed twice, and three additional boundary
cases passed once. Wav2Vec2 hotspot routing passed in all three main trials.
The dispatcher/evaluator evolved between trials; this is not three complete
runs of the final eleven-case suite, a no-skill control, or evidence of
automatic skill discovery. Raw journals and snapshots remain run-local.
Simulated role actions do not certify actual independent role execution.

Real-device rebuild/correctness evidence is separate from these simulations.
Model identities, artifacts and device measurements remain run-local and
are not added to reusable knowledge.
Loading