Skip to content

CondonFM Skills - #4

Merged
ohadmo merged 13 commits into
mainfrom
omosafi/skills
Oct 7, 2026
Merged

ohadmo merged 13 commits into
mainfrom
omosafi/skills

Conversation

@ohadmo

@ohadmo ohadmo commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs.

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo
ohadmo requested a review from caofan September 3, 2026 23:11
@ohadmo ohadmo self-assigned this Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

Thanks for your interest. This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs.

@github-actions github-actions Bot closed this Sep 3, 2026
@ohadmo ohadmo reopened this Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026

Copy link
Copy Markdown

Thanks for your interest. This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs.

@github-actions github-actions Bot closed this Sep 3, 2026
@ohadmo ohadmo reopened this Sep 4, 2026
@ohadmo

ohadmo commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

@ohadmo

ohadmo commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

@greptile-apps

greptile-apps Bot commented Sep 4, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Adds four new skill modules with supporting infrastructure.

The PR appears safe to merge because no blocking failure remains from the previously reported embed evaluation issues.

Summary

The embed evaluation artifacts were regenerated to cover all four current cases, and the benchmark now gives a consistent PASS publication verdict.

  • Updates the benchmark from two to four evaluation tasks, including cases 003 and 004.
  • Signs the current benchmark, evaluation definition, and edge-case fixture.
  • Aligns the benchmark verdict, publication recommendation, and displayed scoring criteria.

Reviews (13) · Last reviewed commit: "Attach NVSkills validation signatures" · Reviewed by Greptile

@ohadmo

ohadmo commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

1 similar comment
@ohadmo

ohadmo commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

@greptileai

@ohadmo

ohadmo commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

1 similar comment
@ohadmo

ohadmo commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

@chrisknvidia

Copy link
Copy Markdown

@ohadmo : I did some debugging of the failures,
Suggestion for the resource-intensive Tier 3 evaluations:
Please add the following to skills/codonfm-score/evals/config.yml

schema_version: 1

harbor:
  timeout_multiplier: 4
  sandbox:
    templates:
      claude-code: harbor-eval-claude-code-8g
      codex: harbor-eval-codex-8g
  • timeout_multiplier: 4 applies to both agents.
  • Each template applies to both the with-skill and baseline arm for that agent.
  • The Claude mapping is especially relevant because the latest remaining failure occurred in the Claude baseline.
  • This provides equal resource headroom, but a rerun is still required.

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 10, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

1 similar comment
@ohadmo

ohadmo commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

@chrisknvidia

Copy link
Copy Markdown

@ohadmo : I tested the two previously failing codonfm-score cases end to end using Claude Opus 5 through NVCARPS → SkillEvaluator → Harbor → Astra.

  • Resolved model: aws/anthropic/bedrock-claude-opus-5
  • Coverage: valid_full
  • Claude results: 4/4 scored, with zero failed attempts or policy refusals
  • Attempts 2–3 were correctly skipped because every case passed on attempt 1
  • Test MR !94
  • Successful Tier 3 job

https://github.com/NVIDIA-BioNeMo/CodonFM/blob/3f95e4e01604dbe7fda83d5563ef7a55d3134772/skills/codonfm-score/evals/evals.json

Could you please add the following per-skill override to skills/codonfm-score/evals/config.yml and rerun the complete PR evaluation?

schema_version: 1

harbor:
  agents:
    claude-code:
      model: aws/anthropic/bedrock-claude-opus-5

This changes only the Claude model for codonfm-score; the Codex configuration remains unchanged. The controlled test confirms that Opus 5 avoids the refusals for the two previously failing cases. The complete PR rerun will confirm it across the full dataset.

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

1 similar comment
@ohadmo

ohadmo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

@ohadmo

ohadmo commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Comment thread skills/codonfm-embed/BENCHMARK.md Outdated
- codonfm-embed-003: malformed sequences_edge_cases.csv (blank split,
  duplicate id, oversized/truncated sequence) — tests whether the agent
  catches known footguns the skill documents explicitly, rather than
  just validating a clean CSV.
- codonfm-embed-004: checkpoint-choice question that requires the
  preprint's benchmark knowledge rather than anything derivable from
  the supplied source code alone — tests the skill's distilled domain
  knowledge, not just code-reading ability.

Both target gaps identified while investigating why codonfm-embed showed
flat/negative lift in Tier 3 runs: the existing two cases are fully
answerable from the shipped source code alone, so a careful baseline can
match a skill-equipped agent on content. These add cases where skill
content should matter more than brute-force source reading.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@ohadmo

ohadmo commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Comment thread skills/codonfm-embed/evals/evals.json
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@ohadmo

ohadmo commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

/nvskills-ci

Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
@ohadmo
ohadmo merged commit 26457c3 into main Oct 7, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants