Skip to content

Add CondonFM skill - #53

Merged
ohadmo merged 1 commit into
mainfrom
omosafi/condonfm-skill
Oct 9, 2026
Merged

ohadmo merged 1 commit into
mainfrom
omosafi/condonfm-skill

Conversation

@ohadmo

@ohadmo ohadmo commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

Signed-off-by: Ohad Mosafi <omosafi@nvidia.com>
@greptile-apps

greptile-apps Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium impact] Adds a new skill for codon sequence analysis.

The PR appears safe to merge, with a non-blocking fix recommended for capped RiboNN sample preparation.

Findings

  1. P2 Discarded rows block sample preparation ▶

Summary

Adds four CodonFM skills for setup, embeddings, variant scoring, and fine-tuning.

  • Registers the family for source sync, plugin inclusion, and skill discovery.
  • Includes matching aggregate copies, evaluation fixtures, reports, and signing artifacts.
  • One non-blocking issue: discarded rows can wrongly stop capped RiboNN sample preparation.
  • Acknowledged by ohadmo in the PR's docs/sync-findings.md: outside-tree links remain broken in catalog copies, and stage_eval_context.py remains upstream maintenance tooling. Both are deferred upstream to preserve published bytes and signatures.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A["CodonFM upstream skills"] --> B["components.d/codonfm.yml"]
  B --> C["open-models-skills/codonfm"]
  C --> D["Plugin inclusion"]
  C --> E["skills.sh.json discovery"]
  D --> F["Aggregate skill copies"]
Loading

Reviews (1) · Last reviewed commit: "Add CondonFM skill" · Reviewed by Greptile

Comment on lines +68 to +73
if sequence in sequence_splits and sequence_splits[sequence] != split:
raise ValueError("Identical CDS appears in different source folds; resolve split leakage before training")
seen_ids.add(row_id)
sequence_splits[sequence] = split
if max_rows_per_split and counts[split] >= max_rows_per_split:
continue

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Discarded rows block sample preparation

prepare() records sequences before checking whether their split is already full. With the default row cap, a discarded training row can make a later validation or test row fail the cross-split check. The helper then produces no CSV, even though the duplicate would not appear in its output.

Check the split cap before checking and recording IDs and sequences. Add a test for this case. The aggregate copy of prepare_ribonn.py needs the same fix.

@ohadmo
ohadmo merged commit 4f1b6a4 into main Oct 9, 2026
4 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant