Repository navigation
CondonFM Skills - #4
Conversation
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
Thanks for your interest. This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs. |
|
Thanks for your interest. This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs. |
|
/nvskills-ci |
|
|
/nvskills-ci |
1 similar comment
|
/nvskills-ci |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
|
/nvskills-ci |
1 similar comment
|
/nvskills-ci |
|
@ohadmo : I did some debugging of the failures,
|
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
1 similar comment
|
/nvskills-ci |
|
@ohadmo : I tested the two previously failing codonfm-score cases end to end using Claude Opus 5 through NVCARPS → SkillEvaluator → Harbor → Astra.
Could you please add the following per-skill override to This changes only the Claude model for codonfm-score; the Codex configuration remains unchanged. The controlled test confirms that Opus 5 avoids the refusals for the two previously failing cases. The complete PR rerun will confirm it across the full dataset. |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
1 similar comment
|
/nvskills-ci |
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
- codonfm-embed-003: malformed sequences_edge_cases.csv (blank split, duplicate id, oversized/truncated sequence) — tests whether the agent catches known footguns the skill documents explicitly, rather than just validating a clean CSV. - codonfm-embed-004: checkpoint-choice question that requires the preprint's benchmark knowledge rather than anything derivable from the supplied source code alone — tests the skill's distilled domain knowledge, not just code-reading ability. Both target gaps identified while investigating why codonfm-embed showed flat/negative lift in Tier 3 runs: the existing two cases are fully answerable from the shipped source code alone, so a careful baseline can match a skill-equipped agent on content. These add cases where skill content should matter more than brute-force source reading. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
/nvskills-ci |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
This repository is read-only and does not accept Pull Requests. Please open an Issue for reproducible bugs.