Skip to content

Add DPO loss chart to LoRA validation docs - #22

Merged
micahtyong merged 2 commits into
mainfrom
dpo-validation-plot
Oct 4, 2026
Merged

micahtyong merged 2 commits into
mainfrom
dpo-validation-plot

Conversation

@micahtyong

@micahtyong micahtyong commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Adds a chart to the "DPO on HHH: Qwen3.5-9B-Base" section of docs/lora_validation.md. It plots DPO loss per step for the cookbook DPO recipe on Spindle and on Tinker, using the same settings (lr 1e-4, β 0.1, batch 256, 10 steps).

DPO loss

Both runs see identical batches (same pair and token counts at every step) and follow the same loss trajectory. Step 9 dpo_loss is 0.6795 on Spindle and 0.6837 on Tinker. The section also links both W&B runs.

🤖 Generated with Claude Code

micahtyong and others added 2 commits October 4, 2026 15:17
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@micahtyong micahtyong changed the title Add DPO loss and reward chart to LoRA validation docs Add DPO loss chart to LoRA validation docs Oct 4, 2026
@micahtyong
micahtyong merged commit e80115d into main Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant