Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 6 additions & 3 deletions ARCHIVE-CUSTODY.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@ TestForge v1.1.7 is one two-skill Augment with several distinct distribution obj
|---|---|---|
| Maintained package | `testforge/` | Current v1.1.7 two-skill source, tools, schemas, examples, evals, adapters, and customer documentation |
| Codex marketplace plugin | `plugins/testforge/` plus `.agents/plugins/marketplace.json` | Repository-native plugin source for `testforge@cd-testforge`; static structure is repository-tested |
| Claude operator upload | `claude-ai/software-verification-v1.1.7.zip` | Current one-skill upload candidate; SHA-256 `0fd9105fdb498259fc0d14aba98907dbcfb0c55b3091e83c68104c47d108e24e` |
| Claude reviewer upload | `claude-ai/verification-reviewer-v1.1.7.zip` | Current one-skill upload candidate; SHA-256 `47399ff6c1a40bbb13db6d63ca597d7d00529527be1c5d59d3d9167e7d5a1ec6` |
| Claude operator upload | `claude-ai/software-verification-v1.1.7.zip` | Current one-skill upload candidate; SHA-256 `7229d1118ae86b6f48bb14cfd69e1a8db48f50b074355f6bead621e1c59d7ac9` |
| Claude reviewer upload | `claude-ai/verification-reviewer-v1.1.7.zip` | Current one-skill upload candidate; SHA-256 `c12f2c5b9753af3cde6c6ad6ec062ea1bca53d97d43fc9ccaa804ca889fba64d` |
| Local v1.1.7 customer kit | `releases/v1.1.7/TestForge-v1.1.7.zip` | Deterministic local candidate; the adjacent `.sha256` file is canonical because this document is itself packaged inside the archive |
| Local v1.1.7 receipts | `releases/v1.1.7/` | Static package, source-parity, and portable archive evidence; no fresh-host activation, customer-outcome, tag, GitHub release, or publication claim |
| Source state | Current `main` contains the v1.1.7 candidate | Tag and GitHub-release state must be established by live remote readback; this retained document does not infer publication from local bytes |
Expand Down Expand Up @@ -38,4 +38,7 @@ There is no retained v1.1.7 portal archive or custody object. The repository-nat
- Record archive name, byte size, SHA-256, member inventory, source revision, and claim boundary in the release receipts for each new object.
- Verify extraction topology and package-relative dependencies before publication.
- After publication, download the public asset and compare it with the governed local object.
- Treat upload, automated scan, review submission, approval, publication, installation, discovery, invocation, and health as separate observed states.
- Treat upload, automated scan, review submission, approval, publication, installation, discovery, invocation, and health as separate observed states.
## Same-version maintenance, 2026-09-05

Source commit `73f650e` moves metered-verification detail behind the existing task trigger while preserving the capacity, reserve, authorization, and response-template safeguards. The current v1.1.7 packages and delivery sidecars are rebuilt from maintained source. The original v1.1.7 tag remains unchanged; replacement asset custody is established separately by readback. This repair carries static package and source-parity evidence, not fresh-host behavioral evidence.
Binary file modified claude-ai/software-verification-v1.1.7.zip
Binary file not shown.
Binary file modified claude-ai/verification-reviewer-v1.1.7.zip
Binary file not shown.
27 changes: 27 additions & 0 deletions delivery/TestForge v1.1.7 Extra.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Description

TestForge is a free two-SKILL verification system for frozen release candidates. `software-verification` reconstructs impact, ranks consequential failure risks, builds meaningful oracles, runs only checks that can change the bounded verdict, and returns a traceable release assessment. `verification-reviewer` then attacks that evidence chain for omissions, weak tests, unsupported claims, and conclusions that outrun the proof. Together they turn verification into a quality ratchet instead of an ornamental pile of green checkmarks.

# Usage Notes

Copy the `.zip` from Additional Files to the chosen harness or Chat project, attach or reference it in chat, and say, `Install this Augment.`

- Give `$software-verification` a completed, frozen candidate and an explicit release-readiness claim. Ordinary implementation work does not need the full TestForge apparatus.
- Keep the builder and verifier roles distinct. A discovered product defect or newly exposed requirement returns to builder custody.
- Use `$verification-reviewer` on the finished evidence package, preferably in a fresh context.
- Every check, retry, artifact, and reviewer pass should be capable of changing the bounded verdict. TestForge does not certify defect freedom or authorize release.

Public GitHub Repo: [TestForge](https://github.com/Stunspot/TestForge)

Project Site: [TestForge verification workbench](https://stunspot.github.io/TestForge/)

# Changelog

2026-09-05 maintenance - Kept hosted-verification safeguards together in their task-specific guide, reducing the instructions loaded for ordinary local verification.

v1.1.7 - Tightened activation to explicit frozen-candidate release verification, enforced decision-changing evidence, and added bounded recovery and stopping rules.
v1.1.6 - Added metered-verification safeguards for quota-limited environments.

# Tags

software verification, release readiness, risk-based testing, adversarial review, evidence, regression gates, behavioral evals, quality assurance, Codex, Claude, Agent SKILL, Augment
18 changes: 18 additions & 0 deletions delivery/TestForge v1.1.7.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# TestForge

Install and onboard the attached **TestForge v1.1.7** Augment on the current AI harness.

The attached ZIP is the supplied customer package. Inspect the archive and its README, QUICK-START, installation guidance, adapters, archive-custody notes, release notes, and package verifier before changing files. The user authorizes installation of this Augment on the current harness only. TestForge contains two coordinated Agent SKILLs: `software-verification` and `verification-reviewer`. Choose the package's documented Codex plugin, standalone skill, Claude, local-shell, or copy-paste path for this host. Keep the two skills and their referenced resources together. Do not install repository tooling, frozen evidence, evaluation results, or maintainer cargo as runtime content unless the documented host path requires it.

Before writing, detect any existing TestForge installation, its location and version, and the applicable install root. Never overwrite, merge, delete, or replace an existing installation unless the package supplies a documented update path and a recoverable rollback is established; otherwise stop and report the collision. If this host supports durable installation, perform the documented installation and any safe required reload. If it cannot install attached Augments, say so plainly and use the packaged fallback or give the exact manual installation path. Do not improvise a different package layout.

Afterward, report these states separately: pre-existing installation and collision result; package integrity when the verifier or checksum is available; installed locations for both skills; rollback or recovery path; host discovery; one explicit invocation of each skill; relevant tool health; and anything not tested. A visible folder is not proof that either skill is active.

Then onboard me:

1. Explain in no more than three sentences that TestForge verifies a frozen release candidate with risk-ranked, decision-changing evidence and independently challenges whether the resulting verdict is supported.
2. State its boundary: TestForge does not prove defect freedom, certify compliance, authorize production access, or replace accountable release judgment.
3. Ask for the frozen candidate, intended release claim, impact surface, constraints, available evidence, and the decision the verdict must support.
4. Offer this first request: "$software-verification Verify this completed frozen candidate for release. Reconstruct what could break, run only decision-changing authorized checks, and give me one evidence-backed assessment."
5. Explain that `$verification-reviewer` should receive the completed evidence package in a fresh context when practical.
6. Begin only when I provide the candidate and boundary. Keep implementation and verification custody distinct; if verification exposes a product defect or new requirement, return it to builder custody rather than silently repairing the candidate.
Binary file added delivery/TestForge.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
22 changes: 2 additions & 20 deletions plugins/testforge/skills/software-verification/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: software-verification
description: "Explicit release-grade adversarial verdict for a frozen software or release candidate; not routine build verification or repair."
description: "☠️ Frozen releases tested for fatal defects."
---

# ☠️ WARNING — ENTER THE CHAPEL PERILOUS
Expand Down Expand Up @@ -51,7 +51,7 @@ Record the target, included and excluded surfaces, constraints, assumptions, kno
Load doctrine at the judgment moment:

- `references/core/risk-based-testing.md` and `test-layer-selection.md` for prioritization and the smallest credible evidence set.
- `references/core/metered-verification.md` before proposing or invoking hosted CI, device/browser farms, paid cloud tests, or any other quota-limited verification.
- `references/core/metered-verification.md` before proposing or invoking hosted CI, device/browser farms, paid cloud tests, or any other quota-limited verification; follow its mandatory capacity, usage, reserve, authorization, and response-template contract before dispatch.
- `references/core/oracle-design.md`, `boundary-and-equivalence.md`, and `state-transition-testing.md` for discriminating assertions and scenario design.
- `references/core/test-smells.md` for mock boundaries and deceptive tests.
- `references/core/release-assessment.md` for release status.
Expand All @@ -70,24 +70,6 @@ For each scenario, state preconditions, action, expected observations, forbidden

Create or repair repository-compatible tests, fixtures, builders, commands, and records. Production-code changes, dependency installation, weakened or deleted tests, material snapshot updates, CI/deployment edits, destructive operations, production targets, active security checks, and external publication require explicit human authority at the point of action.

## Preflight metered verification

Before recommending or invoking a quota-limited verification service, obtain a current capacity snapshot from an authoritative provider API, provider UI, or identified operator observation. Record the provider, observation time, capacity state, remaining allowance when observable, refresh or billing-cycle boundary, paid-overage state, principal-set reserve, and the evidence source. Missing access to the allowance is `unknown`, never zero and never permission to probe by launching a job.

Estimate the complete planned consumption before execution. Include every trigger, matrix expansion, job, retry or rerun allowance, runner ceiling, and applicable provider billing multiplier. Do not launch a metered check merely to discover whether capacity exists. Run `scripts/assess_metered_verification.py` against the recorded snapshot and plan; a hold result blocks automatic invocation.

Use provider-hosted execution only when the provider boundary is itself under test or an already-authorized acceptance contract requires it. Otherwise prefer the smallest credible local, clean-host, self-hosted, or batched substitute and state the exact guarantee the substitution does not establish. Avoid duplicate push-and-pull-request execution unless each trigger supplies decision-relevant evidence. Paid overage never becomes authorized merely because it is technically available, and the assessor never grants or authenticates spend authority.

In the response, state the capacity classification and dispatch decision before any command. Even when allowance or a current multiplier is unknown, expand every known trigger, matrix job, attempt, and ceiling. Write the arithmetic and raw runner-minute total explicitly, then identify the missing multiplier rather than dropping the fan-out. On every hold, name at least one credible substitute and the exact hosted-provider guarantee it would leave unproven—for hosted CI, normally provider runner/image behavior and the provider's own trigger, matrix, permission, secret, artifact, and status integration. Never invent a `paid_overage_authorization` field, override flag, dispatch command, or other route by which caller-authored text could impersonate the human decision. Stop at a bounded authority request that names the exact run, maximum paid minutes, maximum monetary spend when price data is available, expiry, and billing scope; the human's later answer must still be resolved by a trusted dispatcher outside the assessor.

Keep every metered preflight short and decision-shaped. Use these five headings exactly once: `Capacity`, `Expansion`, `Decision`, `Substitute`, and `Authority`. Under `Expansion`, write one complete equation: `triggers × matrix jobs × attempts × ceiling minutes × provider multiplier = estimated billed minutes`. When the current multiplier is unobserved, mark it explicitly `unknown` and separately state the raw runner-minute total through the ceiling term. Never label the intermediate job-attempt count as runner-minutes. `Substitute` is mandatory on every hold and must pair the proposed route with a direct sentence beginning `This substitute does not prove:` followed by the provider runner/image, trigger/matrix, permission/secret, artifact, and status-integration guarantees that remain absent from the acceptance claim. A missing local host or command does not excuse omitting the route: describe a local, clean-host, self-hosted, or batched substitute generically as `PREPARED — NOT EXECUTED` and state what capability would execute it. Do not invent a local command or file path; use a repository-documented route only when observed. Do not narrate internal debate or repeat corrected calculations; provide the final conservative arithmetic and decision.

Load `assets/templates/metered-verification-response.md` and complete it from the observed case. It is the response contract, not an optional example.

Copy snapshot facts exactly; do not replace a supplied remaining-validity interval, observation, refresh boundary, reserve, or multiplier with a guessed timestamp or default. Always report `required_with_reserve_minutes = estimated_minutes + reserve_minutes`. If paid capacity is available but unauthorized, report `included_available_after_reserve = max(remaining_minutes - reserve_minutes, 0)` and `maximum_paid_minutes_required = max(estimated_minutes - included_available_after_reserve, 0)`. The bounded human request uses that single maximum, never a range or “if reserve logic dictates” alternative. Example: a 45-minute plan, 15 included minutes, and a 10-minute reserve require 55 minutes with reserve, leave 5 included minutes usable, and require at most 40 paid minutes.

Reserve is retained, not spendable capacity. Calculate `estimated_minutes` from the jobs, then `required_with_reserve_minutes = estimated_minutes + reserve_minutes`. For example, 15 remaining minutes, a 10-minute reserve, and a 45-minute plan means 55 minutes are required to run while retaining the reserve; it does not mean 25 non-paid minutes are available.

For authorization denials, observe protected post-state, downstream effects, secret-bearing output, and audit behavior where the contract supplies it; status alone is not the oracle. If active security scope is unauthorized, stop the active action but preserve a safe plan and name the complete re-entry packet: accountable owner permission, target and environment, time window, rate and concurrency bounds, prohibited actions, data-handling rules, and stop contact.

## Validate what is exact; interpret what remains semantic
Expand Down
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
interface:
display_name: "TestForge Verification Operator"
short_description: "Judge a frozen release candidate"
short_description: "☠️ Frozen releases tested for fatal defects."
default_prompt: "Use $software-verification to attack this frozen candidate with only decision-changing checks, then issue one bounded release verdict."
Loading