Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions .github/workflows/request-nvskills-ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
name: Request NVSkills CI

on:
issue_comment:
types: [created]
pull_request:
types: [opened, reopened, synchronize, ready_for_review]
push:

jobs:
request:
if: >
github.event_name == 'pull_request' ||
(github.event_name == 'issue_comment' &&
github.event.issue.pull_request &&
startsWith(github.event.comment.body, '/nvskills-ci')) ||
(github.event_name == 'push' &&
github.actor == (vars.NVSKILLS_SIGNATURE_PUSH_ACTOR || 'nv-skills-ci[bot]') &&
startsWith(github.event.head_commit.message, vars.NVSKILLS_SIGNATURE_COMMIT_TITLE || 'Attach NVSkills validation signatures'))
permissions:
contents: read
pull-requests: read
statuses: read
uses: NVIDIA/skills/.github/workflows/team-request.yml@main
secrets:
NVSKILLS_CI_DISPATCH_TOKEN: ${{ secrets.NVSKILLS_CI_DISPATCH_TOKEN }}
6 changes: 5 additions & 1 deletion requirements.txt
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
# xFormers CUDA wheels are published on the PyTorch index.
--extra-index-url https://download.pytorch.org/whl/cu124

# --------- pytorch --------- #
torch==2.5.1
torchvision==0.20.1
Expand All @@ -22,9 +25,10 @@ pre-commit==4.0.1 # hooks for applying linters on commit
rich==13.9.4 # beautiful text formatting in terminal
pytest==8.1.1 # tests
sh==2.2.2 # for running bash commands in some tests (linux/macos only)
python-dotenv==1.0.1
transformers==4.54.1
polars==1.12.0
xformers==0.0.28.post3 --index-url https://download.pytorch.org/whl/cu124
xformers==0.0.28.post3
ninja==1.11.1.1
einops==0.8.0
ipython-autotime==0.3.2
Expand Down
122 changes: 122 additions & 0 deletions skills/codonfm-embed/BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,122 @@
# Skill Benchmark: codonfm-embed

> ✅ **Overall verdict: PASS — Recommended for publication**
## Publication Recommendation

Recommended for publication based on the completed evaluation evidence in this report.

## Evaluation Metadata

- Skill: `codonfm-embed`
- Evaluation date: 2026-10-07
- Evaluator version: `1.5.6`
- Agents: Claude Code (`aws/anthropic/bedrock-claude-opus-4-8`), Codex (`openai/openai/gpt-5.5`)
- Tasks: 4 evaluation tasks (4 positive)
- Dataset digest: `sha256:8ef619d6a9074d8b7fdc6224f220bb79aad6e37e2222f3290580290b7a6afbeb` (skill-evaluator-dataset-snapshot/1)
- Attempts per task: 1
- Environment: `k8s-sandbox`
- Tier 2 evidence: required for publication
- Tier 3 evidence: required for publication

Each task attempt ran in its own isolated sandbox pod.

## What This Report Answers

The three-tier evaluation checks whether the skill:

- is safe to use;
- produces correct answers;
- is discovered and activated when needed;
- helps the agent complete the user's goal and expected workflow; and
- avoids wasted skill and tool usage.

## Results at a Glance

| Measure | Claude Code (Baseline → Skill Uplift) | Codex (Baseline → Skill Uplift) |
|---|---:|---:|
| Overall | 96.1% — baseline ran, but no comparable score was available; uplift unavailable | 93.4% — baseline ran, but no comparable score was available; uplift unavailable |
| Security | 75.0% → 100.0% (+25.0 points) | 75.0% → 100.0% (+25.0 points) |
| Correctness | 100.0% → 100.0% (±0.0 points) | 100.0% → 100.0% (±0.0 points) |
| Discoverability | 96.3% — baseline ran, but no comparable score was available; uplift unavailable | 86.3% — baseline ran, but no comparable score was available; uplift unavailable |
| Effectiveness | 85.6% → 90.6% (+5.0 points) | 98.8% → 93.8% (-5.0 points) |
| Efficiency | 93.6% — baseline ran, but no comparable score was available; uplift unavailable | 86.9% — baseline ran, but no comparable score was available; uplift unavailable |

**How to read this table:** baseline is the same task attempted without the target skill. Scores are rounded to one decimal; threshold-adjacent values use additional precision so their displayed band matches the verdict. Uplift is derived from those displayed scores and shown in percentage points.

Example: `47.0% → 92.0% (+45.0 points)` means the skill-assisted run scored 92.0%, 45.0 percentage points above its 47.0% no-skill baseline.

## Token Usage

Actual Tier 3 execution usage is reported for every observed agent/case pair and both conditions.

| Agent | Dataset case | With skill | Without skill | Delta | Change | Coverage |
|---|---|---:|---:|---:|---:|---|
| claude-code | All cases | 1,152,349 | 5,695,914 | -4,543,565 | -79.77% | skill 4/4; base 4/4 |
| claude-code | codonfm-embed-001 | 544,672 | 1,934,434 | -1,389,762 | -71.84% | skill 1/1; base 1/1 |
| claude-code | codonfm-embed-002 | 243,222 | 950,364 | -707,142 | -74.41% | skill 1/1; base 1/1 |
| claude-code | codonfm-embed-003 | 265,873 | 2,038,952 | -1,773,079 | -86.96% | skill 1/1; base 1/1 |
| claude-code | codonfm-embed-004 | 98,582 | 772,164 | -673,582 | -87.23% | skill 1/1; base 1/1 |
| codex | All cases | 487,427 | 1,838,975 | -1,351,548 | -73.49% | skill 4/4; base 4/4 |
| codex | codonfm-embed-001 | 106,660 | 518,881 | -412,221 | -79.44% | skill 1/1; base 1/1 |
| codex | codonfm-embed-002 | 159,116 | 322,872 | -163,756 | -50.72% | skill 1/1; base 1/1 |
| codex | codonfm-embed-003 | 157,215 | 924,873 | -767,658 | -83.00% | skill 1/1; base 1/1 |
| codex | codonfm-embed-004 | 64,436 | 72,349 | -7,913 | -10.94% | skill 1/1; base 1/1 |
| ALL AGENTS | Dataset aggregate | 1,639,776 | 7,534,889 | -5,895,113 | -78.24% | skill 8/8; base 8/8 |

Prompt tokens include cached reads, so total tokens are `prompt + completion` (cached is not added twice). The Efficiency score uses `(prompt - cached) + completion`. N/A means the relevant trajectory counters were not available; coverage is never estimated.

## Tier Status

| Tier | Purpose | Status | Evidence |
|---|---|---|---|
| Tier 1 | Static validation | **PASSED WITH OBSERVATIONS** | 11 validator(s); 3 finding(s) |
| Tier 2 | Semantic deduplication | **PASSED** | 2 validator(s); 0 finding(s) |
| Tier 3 | Live agent evaluation | **PASS** | 2 agent(s); 4 task(s) |

## Findings and Observations

<details>
<summary>Show detailed findings and successful checks</summary>

- **MEDIUM** QUALITY/quality_correctness: Instructions don't mention 'run_script' (`skills/codonfm-embed/SKILL.md`)
- **MEDIUM** SECURITY/Unknown (LP3): MCP Least Privilege: Without declared permissions the skill's intent is opaque and cannot be validated. (`SKILL.md:1`)
- **LOW** SCRIPT_LINT/magic_numbers: validate_inputs.py contains magic numbers (`skills/codonfm-embed/scripts/validate_inputs.py`)

</details>

## Scoring Methodology

<details>
<summary>Show dimension definitions, source signals, and thresholds</summary>

| Dimension | Question | Scored signals |
|---|---|---|
| Security | Is it safe to use? | `security` (100%) |
| Correctness | Is the answer correct? | `accuracy` (100%) |
| Discoverability | Was the right skill loaded when needed? | `skill_execution` (100%) |
| Effectiveness | Did the skill help complete the task? | `goal_accuracy` (50%) + `behavior_check` (50%) |
| Efficiency | Did it avoid wasted tool calls and token usage? | `skill_efficiency` (50%) + `token_efficiency` (50%) |

- Dimension bands: PASS at 50% or above; NEUTRAL from 40% to below 50%; FAIL below 40%.
- Overall Tier 3 lift: PASS at +5 points or more; FAIL at -10 points or less; values between those bands are NEUTRAL.
- Overall verdict: PASS only when every configured dimension passes for at least one supported agent. Lift is reported as diagnostic evidence and does not override this gate.
- The 50% attempt pass threshold is a separate per-task gate; it is not the dimension pass threshold.
- Effectiveness is the equal-weight mean of goal completion (`goal_accuracy`) and expected workflow adherence (`behavior_check`).
- Efficiency is 50% tool-call productivity (the backward-compatible `skill_efficiency` wire id) and 50% `token_efficiency`. Positive-case skill routing is scored under Discoverability, not Efficiency; a negative case without a routing target is N/A. N/A sources are omitted, remaining weights are renormalized, and the dimension is marked partial.

Signals present in this run:

- `security` (Security): unsafe operations, secret leakage, and unauthorized access.
- `skill_execution` (Skill Execution): whether the expected skill was selected, decoys were avoided, and the workflow executed.
- `skill_efficiency` (Tool Productivity): tool-call productivity (legacy wire id; routing is scored under Discoverability).
- `accuracy` (Accuracy): final-answer correctness against the reference answer.
- `goal_accuracy` (Goal Accuracy): whether the user's goal was achieved.
- `behavior_check` (Behavior Check): whether the expected workflow behavior was followed.
- `token_efficiency` (Token Efficiency): actual uncached prompt plus completion usage (50% of Efficiency).

</details>

## Freshness

Regenerate this benchmark when the skill, evaluation dataset, target agent/model, evaluator version, environment, or scoring policy changes.
181 changes: 181 additions & 0 deletions skills/codonfm-embed/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,181 @@
---
name: codonfm-embed
description: Validate coding-sequence CSVs, extract public CodonFM/Encodon embeddings, and choose checkpoints for downstream property modeling.
license: Apache-2.0
metadata:
author: "NVIDIA BioNeMo <bionemofeedback@nvidia.com>"
tags: [biology, codonfm, embeddings]
---

# Extract public Encodon embeddings

## Purpose

Extract one frozen CLS vector per coding sequence with public Encodon v1.
Support input validation, command preparation, extraction, and checkpoint
selection for translation efficiency, expression, or mRNA stability modeling.
Extraction does not automatically train a downstream regressor.

## Prerequisites

- Validation needs Python 3 standard library only; no GPU, weights, or API key.
- Execution needs the public CodonFM checkout, its `requirements.txt` environment,
a compatible NVIDIA GPU, and local checkpoint weights. A metadata JSON is not
a checkpoint. A `.safetensors` file needs its sibling `config.json`; `.ckpt`
checkpoints are also supported by the public loader.
- Run `python -m src.runner` from the CodonFM repository root. In an isolated
workspace, use supplied source artifacts; source paths below are relative to
that checkout or source archive, not this skill directory.

## Inputs

Input source precedence: explicit user prompt arguments, then supplied
files/checkpoint metadata, then inspected public runner defaults. Resolve
conflicting model names and checkpoint metadata before execution. Supplied 80M
metadata is useful for preparing an 80M command; it does not restrict an
open-ended recommendation to that size.

Required for validation: a CSV. Required for extraction: the CSV, checkpoint,
matching model name, and output directory. Optional: context length and batch
size overrides. Checkpoint-selection questions can be answered without a CSV.

| Input | Requirement or default |
| --- | --- |
| Sequence CSV | Columns `id`, `ref_seq`, `value`, `split`; extra columns allowed |
| `id` | Nonblank, unique IDs for unambiguous output association |
| `ref_seq` | Coding sequence, uppercase DNA `A/C/G/T`, length divisible by three; public dataset converts uppercase `U` to `T` |
| `value` | Numeric label; use `0.0` for new extraction-only data, preserve supplied labels |
| `split` | Only exact `test` values enter extraction; blank/other values are excluded |
| Checkpoint and model | Match weights/config to `encodon_80m`, `encodon_600m`, or `encodon_1b` |
| Context length | Public runner default `2048` tokens, including CLS and SEP |
| Output directory | A fresh run directory with an empty predictions directory |

## Instructions

1. **Choose the requested workflow.** For a checkpoint/performance question,
read [checkpoint selection](references/checkpoint-selection.md) and answer
from public benchmark evidence. For the strongest published downstream
results, prefer the public **1B random-mask checkpoint** when resources allow;
80M is a demonstration or resource-constrained choice. A small labeled set
alone does not establish that 80M frozen features are better. Do not download
weights or inspect the entire source tree just to make a recommendation.
2. **Inspect supplied source only where needed.** Confirm runner/config,
`src/data/codon_bert_dataset.py`, `src/data/preprocess/codon_sequence.py`,
`src/inference/encodon.py`, or `src/utils/pred_writer.py` for the relevant
behavior. Read ZIP members with `zipfile.ZipFile.namelist()` and `.read()`;
source inspection does not need extraction. If a checkout is needed, use a
new directory from `tempfile.mkdtemp()` or `mktemp -d`, without deleting or
overwriting an existing directory. For Decodon support questions, inspect
runner/config and model/inference modules, cite the inspected files, explain
the missing public implementation, and finish there.
3. **Validate the CSV before running extraction.** Run the bundled checker
below with the intended context length. Report per-row verdicts using CSV
row numbers as well as IDs, since IDs can repeat. Separate excluded rows,
invalid inputs, duplicate-ID warnings, and truncation. Propose fixes without
silently rewriting supplied data. The checker is a preflight, not model
execution or proof of biological CDS validity.
4. **Deliver the requested preparation or execution.** For preparation, return
a complete command with resolved paths (or clearly identified prerequisites),
the test-row count, validation findings, and the output contract below.
Include all task/dataset/process flags in the final answer, even if already
shown in a tool call. For extraction, reuse/download the chosen checkpoint
when needed, execute once resources are ready, and verify the saved arrays.
If resources are missing, finish preparation and state what is missing.

## Available Scripts

| Script | Purpose | Arguments |
| --- | --- | --- |
| [validate_inputs.py](scripts/validate_inputs.py) | Read-only CSV validation and per-row verdicts | Required CSV path; optional `--context-length` (default `2048`) |

Run the preflight with Python; `CODONFM_SKILL_DIR` is the directory containing this file:

```bash
python "$CODONFM_SKILL_DIR/scripts/validate_inputs.py" "$CODONFM_DATA_PATH" \
--context-length 2048
```

The checker prints JSON. Exit `0` means no findings, `1` means row findings to
review (including exclusions/warnings), and `2` means a file/schema error.
Neither warnings nor exclusions imply that the public runner will crash.

## Output Format

The checker emits JSON with `total_rows`, `test_rows`, `excluded_rows`,
`context_length`, `codon_limit`, `warnings`, and `rows`. Each row records its
one-based data-row number (excluding the header), ID, split, verdict, issues,
sequence/value validity, and retained/lost codons. A file/schema error emits
`error` and `csv` instead. These are preflight findings, not generated embeddings.

## Examples

Set `CODONFM_DATA_PATH` to the CSV, `CODONFM_CHECKPOINT_PATH` to the weights,
`CODONFM_MODEL_NAME` to the matching architecture, and `CODONFM_RUN_DIR` to a
fresh output directory. Substitute actual paths in a prepared command:

```bash
python -m src.runner eval \
--task_type embedding_prediction \
--process_item codon_sequence \
--dataset_name CodonBertDataset \
--exp_name embed_extract \
--model_name "$CODONFM_MODEL_NAME" \
--checkpoint_path "$CODONFM_CHECKPOINT_PATH" \
--data_path "$CODONFM_DATA_PATH" \
--context_length 2048 \
--num_nodes 1 \
--num_gpus 1 \
--num_workers 0 \
--val_batch_size 2 \
--out_dir "$CODONFM_RUN_DIR" \
--predictions_output_dir "$CODONFM_RUN_DIR/predictions"
```

For a low-cost demonstration, `encodon_80m` matches
`nvidia/NV-CodonFM-Encodon-80M-v1`, revision
`399ca9fe17b57941a7bebc6788033919b417413c`, file
`NV-CodonFM-Encodon-80M-v1.safetensors` and sibling `config.json`.

## Outputs

- Under `--predictions_output_dir`, `embeddings_merged.npy` contains frozen
final-layer CLS vectors, shape `(processed_rows, hidden_size)`.
- `ids_merged.npy` is index-aligned: embedding row `i` belongs to ID row `i`.
Use these IDs to join to the CSV; do not assume every CSV row was retained.
Duplicate IDs make that join ambiguous even when extraction succeeds.
- For the one-GPU example, verify both arrays have the expected test-row count,
embeddings are finite, and width matches checkpoint config (`1024` for 80M,
`2048` for 600M/1B). Do not fabricate arrays for a preparation-only request.

The public checkout's downstream-model references are:

- `notebooks/4-EnCodon-Downstream-Task-riboNN.ipynb`
- `notebooks/5-EnCodon-Downstream-Task-mRFP-expression.ipynb`
- `notebooks/6-EnCodon-Downstream-Task-mRNA-stability.ipynb`

## Limitations

- Public v1 has no Decodon model/inference implementation, Decodon notebooks,
`notebooks/te_predictor.py`, or `notebooks/mfe_predictor.py`.
- At context length `2048`, retain the first `2046` codons; any remaining
3-prime sequence is lost. Increasing the flag does not validate a longer
context. Disclose deliberate cropping or a separate chunking/aggregation
strategy; neither is equivalent to embedding the complete sequence once.
- `--dryrun` builds runtime configuration, may create directories, and needs
ML dependencies; it reads neither the CSV nor the weights and is not input
validation.
- Do not claim a benchmark-trained regressor generalizes to a new organism,
cell type, or assay without new labeled validation data.
- Do not invoke this skill for a generic expression-prediction request that
does not mention CodonFM or Encodon.

## Troubleshooting

| Symptom | Cause and action |
| --- | --- |
| Missing `split` column | Eval requests the test split despite the dataset docstring calling this column optional; add an explicit split column to a corrected copy |
| Fewer output rows | Blank/non-`test` split values are silently filtered; set intended extraction rows to exact `test` in a corrected copy |
| Repeated output IDs | Duplicate input IDs are not rejected; assign unique IDs while preserving a mapping to the original rows |
| Oversized sequence | Preprocessing truncates at `context_length - 2` codons; report retained/lost lengths and agree on a sequence-handling strategy |
| Missing weights or dependencies | Complete validation/command preparation; metadata and `--dryrun` do not substitute for weights |
| Merge failure on a repeated run | The writer scans `.npy` files; use a fresh predictions directory to avoid stale shards or merged arrays |
4 changes: 4 additions & 0 deletions skills/codonfm-embed/agents/openai.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
interface:
display_name: "CodonFM Embeddings"
short_description: "Extract public Encodon sequence embeddings"
default_prompt: "Use $codonfm-embed to extract Encodon embeddings from my coding-sequence CSV."
Loading
Loading