Repository navigation
Conversation
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
|
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
| ## Evaluation Results: <br> | ||
| | Measure | Claude Code (Baseline → Skill Uplift) | Codex (Baseline → Skill Uplift) | | ||
| |---|---:|---:| | ||
| | Overall | 97.1% | 92.6% | |
There was a problem hiding this comment.
Neutral verdict missing from card
The refreshed card lists strong evaluation scores but omits the benchmark’s new NEUTRAL publication verdict and its recommendation to collect more evidence before deciding to publish. Someone using the card as a release summary could mistake the scores for a positive publication recommendation. Please include the verdict or link to the benchmark alongside the results.
Knowledge Base Used: Skill quality and compliance controls
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
9daa040 to
1494369
Compare
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
|
/nvskills-ci |
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
e5c239f to
159fa23
Compare
|
/nvskills-ci |
| result["issues"].append("Parabricks will generate read groups. Confirm sample metadata if read-group identity matters; no metadata was inferred by this helper.") | ||
| result["notes"].append("Draft only: paths, reference/index compatibility, numeric ranges, runtime readiness, and biological output parity are not validated.") | ||
| result.update(status="needs_review" if result["issues"] else "draft", | ||
| draft_argv=argv, command=shlex.join(argv)) |
There was a problem hiding this comment.
Translator fails on Python 3.7
If the host uses Python 3.7, every otherwise valid translation reaches shlex.join, which is unavailable in that version. The helper then raises an uncaught AttributeError instead of producing a command, even though its reference says only that Python 3 is required. Specify Python 3.8 or newer for this helper, or use a compatible command formatter.
Signed-off-by: nvskills-svc-account <svc-nvskills-signing@nvidia.com>
| @@ -0,0 +1,153 @@ | |||
| # Skill Benchmark: parabricks | |||
|
|
|||
| > **Overall verdict: NEUTRAL — One or more dimensions remain below PASS** | |||
There was a problem hiding this comment.
Neutral verdict cites missing decline
The refreshed benchmark says its NEUTRAL verdict is because one or more dimensions are below PASS. But every reported Codex dimension exceeds the report’s 50% PASS threshold, and every Claude Code score is unavailable. This explanation points readers to a low score the results do not show, making it harder to tell whether the limitation is missing evaluation evidence instead.
Knowledge Base Used: Skill quality and compliance controls
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
- Update tool references and command conventions for Parabricks 4.7.1 - Improve planning triggers, option translation, and input safeguards - Replace the Python readiness helper and adapt tests and eval fixtures Signed-off-by: Angel Pizarro <apizarro@nvidia.com>
…s-skill Signed-off-by: Angel Pizarro <apizarro@nvidia.com>
| @@ -0,0 +1,218 @@ | |||
| --- | |||
There was a problem hiding this comment.
The new parabricks/parabricks/ directory exists in the generated plugin payload but not in the source skill. The plugin freshness check treats its files as extra content and fails. The top-level distributed SKILL.md also remains at version 1.2.1 while the source is 1.2.4, so users receive older guidance. Regenerate the flat payload from the source instead of adding a second skill directory inside it.
Knowledge Base Used: Plugin synchronization and generated skills
| else | ||
| "$@" >"$TMP_DIR/out" 2>"$TMP_DIR/err" </dev/null | ||
| CMD_RC=$? |
There was a problem hiding this comment.
If neither timeout nor gtimeout is installed, this fallback runs external probes without a time limit. A stalled Docker command can therefore leave the readiness check hanging indefinitely even when the caller supplied --timeout, rather than returning a report.
Knowledge Base Used: Genomic acceleration library integrations
| - Tasks: 18 evaluation tasks (17 positive, 1 negative) | ||
| - Dataset digest: `sha256:6e9efe466a1409e4cbcd58924a921028b12d3af7cce343c86b975cfbf78f9f1b` (skill-evaluator-dataset-snapshot/1) |
There was a problem hiding this comment.
The refreshed results cover seven live tasks, but this metadata still says 18 tasks and gives a different dataset digest from the seven-task skill card. Readers cannot tell which dataset produced the published scores and verdict, making the release evidence harder to assess.
Knowledge Base Used: Skill quality and compliance controls
Comments Outside DiffThese findings could not be posted inline.
|
No description provided.