Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ Spindle uses three separate credentials:
Generate an API key for your Spindle server and store it as a [Modal Secret](https://modal.com/docs/guide/secrets):

```bash
export TINKER_API_KEY="$(openssl rand -hex 32)"
export TINKER_API_KEY="tml-$(openssl rand -hex 32)"
modal secret create spindle-api TINKER_API_KEY="$TINKER_API_KEY"
```

Expand Down
33 changes: 33 additions & 0 deletions docs/lora_validation.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,3 +94,36 @@ samples, lr 1e-4, PPO clip 0.8/1.28, no KL. Both runs use an 8×H200 trainer and

The Spindle curve covers the steps completed at the time of writing; the native
Miles curve is the full run.

## DPO on HHH: Qwen3.5-9B-Base (2026-10-02)

DPO runs on the stock `qwen35-9b-lora-16k` deployment with no extra config. The
tinker-cookbook DPO recipe uses `forward_backward_custom` (a `forward` followed by
`cross_entropy` with per-token weights) and gets reference logprobs from
`compute_logprobs` on a sampler published from the step-0 weights. The run is the
[cookbook DPO recipe](https://github.com/thinking-machines-lab/tinker-cookbook/blob/main/tinker_cookbook/recipes/preference/dpo/README.md)
pointed at a Spindle server, with lr 1e-4 for 10 steps:

```bash
export TINKER_BASE_URL=https://your-modal-server-url
export TINKER_API_KEY=...
# dataset=hhh: https://huggingface.co/datasets/Anthropic/hh-rlhf
python -m tinker_cookbook.recipes.preference.dpo.train \
base_url=$TINKER_BASE_URL model_name=Qwen/Qwen3.5-9B-Base dataset=hhh \
renderer_name=role_colon learning_rate=1e-4 dpo_beta=0.1 max_steps=10 \
log_path=/tmp/dpo-hhh-experiment
```

| Metric | Step 0 | Step 9 |
| --- | --- | --- |
| dpo_loss | 0.6939 | 0.6795 |
| accuracy | 0.480 | 0.557 |
| margin | −0.0011 | 0.0502 |
| chosen reward | −0.0001 | 0.0855 |
| rejected reward | 0.0011 | 0.0353 |

- Margin and accuracy rise over the run, with chosen rewards pulling away from
rejected ones.
- A warm step takes about 21 s, plus 9–14 s to compute reference logprobs on
the sampler pool.
- [W&B run](https://wandb.ai/modal-labs/spindle-dpo-validation/runs/0z0obns5).
7 changes: 6 additions & 1 deletion src/spindle/inference/sampling.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,12 @@
RETRY_MAX_DELAY_SECONDS = 5.0
TCP_KEEPALIVE_OPTIONS = (
(socket.SOL_SOCKET, socket.SO_KEEPALIVE, 1),
(socket.IPPROTO_TCP, socket.TCP_KEEPIDLE, 60),
# TCP configs validated on MacOS and Linux.
(
socket.IPPROTO_TCP,
getattr(socket, "TCP_KEEPIDLE", None) or socket.TCP_KEEPALIVE,
60,
),
(socket.IPPROTO_TCP, socket.TCP_KEEPINTVL, 60),
(socket.IPPROTO_TCP, socket.TCP_KEEPCNT, 5),
)
Expand Down
Loading