Spindle is an open-source Tinker-compatible API with support for
multi-tenant LoRA training and
single-tenant full-parameter training.
Trainers run forward_backward and optim_step calls, then publish updated weights to autoscaling
sampling replicas managed by the Stitch protocol.
Owning your training stack is incredibly valuable, especially if you don't have to manage the underlying infrastructure.
Out of the box, we provide defaults that work for the vast majority of use cases. But when you need to incorporate custom forks into the backend/trainer layer, tweak compute parameters for optimal price and performance, or even control the entire training/scheduling runtime, you have the power to do so.
Below, we detail how to setup your own Spindle server. See here for a complete example.
- Multi-tenant LoRA training. Several LoRA adapters share a base model on a Miles/Megatron backend, each with its own gradients and optimizer state. This path works with the Tinker SDK out of the box via
create_lora_training_client. See Working with Multi-LoRA. - Single-tenant full-parameter training (FFT). Each run gets dedicated training containers in your Modal workspace. Tinker (v0.25) doesn't natively support FFT, so we added support ourselves using a few Spindle helpers. See Working with Full Fine-Tunes for more.
Spindle requires Python 3.12. With uv, create a project and install the library:
uv init --python 3.12 my-spindle-project
cd my-spindle-project
uv add 'modal-spindle @ git+https://github.com/modal-projects/spindle.git'If someone has already deployed Spindle for you, skip to Run a Tinker script with the server URL and Spindle API key they provide. API clients don't need Modal credentials or proxy tokens.
Authenticate with a Modal account that can deploy apps and create secrets, and select the environment to deploy into before creating secrets, so the secrets and the deployment end up in the same environment:
modal token new
export MODAL_ENVIRONMENT=your-environment
modal environment listFor automation, supply existing MODAL_TOKEN_ID / MODAL_TOKEN_SECRET credentials instead of the interactive login.
Spindle uses three separate credentials:
| Credential | Purpose | Who needs it |
|---|---|---|
| Modal API token / local profile | Manage Modal resources | Deployer |
TINKER_API_KEY in the spindle-api secret |
Authenticate calls to the Spindle API | Deployer and API clients |
Proxy token in the spindle-proxy secret |
Let Spindle reach protected sampler pools | Deployed control plane and trainers |
Generate an API key for your Spindle server and store it as a Modal Secret:
export TINKER_API_KEY="tml-$(openssl rand -hex 32)"
modal secret create spindle-api TINKER_API_KEY="$TINKER_API_KEY"The Spindle server requires a Modal Proxy Token.
Create one, export the printed Modal-Key and Modal-Secret, and allow it in your environment:
modal workspace proxy-tokens create --name spindle
export MODAL_PROXY_TOKEN_ID='<Modal-Key>'
export MODAL_PROXY_TOKEN_SECRET='<Modal-Secret>'
modal workspace proxy-tokens allow "$MODAL_PROXY_TOKEN_ID" "$MODAL_ENVIRONMENT"If you deploy with service-user credentials or lack permission to create workspace proxy tokens, have a workspace owner or manager provision an allowed token and export that pair instead.
Store this token under the spindle-proxy secret in the same environment:
modal secret create spindle-proxy \
MODAL_PROXY_TOKEN_ID="$MODAL_PROXY_TOKEN_ID" \
MODAL_PROXY_TOKEN_SECRET="$MODAL_PROXY_TOKEN_SECRET"To deploy the Spindle server, simply run:
spindle deployYou'll want to save the server URL from the above deployment output.
A single deployment can serve multiple concurrent independent training jobs across different base models, training parameterizations, and experiment scales. GPUs are allocated on demand upon the first training/inference requests, so the first step may incur longer cold start and model compilation times. See cold starts and capacity configuration before running a larger workload.
Any Tinker-compatible script can be run out of the box against a Spindle server just by changing the base URL and API key.
First, set these two variables from above:
export TINKER_BASE_URL='https://your-server-url.modal.run'
export TINKER_API_KEY='your-spindle-api-key'As an example, the following script implements a single step RL update, which exercises the full generation/training/sampler publication path.
import os
import tinker
from tinker import types
service = tinker.ServiceClient(
base_url=os.environ["TINKER_BASE_URL"],
api_key=os.environ["TINKER_API_KEY"],
)
training = service.create_lora_training_client(
base_model="Qwen/Qwen3.5-9B-Base",
rank=16,
)
tokenizer = training.get_tokenizer()
prompt_tokens = tokenizer.encode(
"What is 2 + 2? Answer with only the number.",
add_special_tokens=True,
)
prompt = types.ModelInput.from_ints(prompt_tokens)
params = types.SamplingParams(max_tokens=16, temperature=1.0)
sampling = training.save_weights_and_get_sampling_client()
response = (
sampling.sample(
prompt=prompt,
num_samples=1,
sampling_params=params,
)
.result(timeout=3600)
.sequences[0]
)
answer_tokens = list(response.tokens)
logprobs = list(response.logprobs or [])
assert answer_tokens and len(logprobs) == len(answer_tokens)
answer = tokenizer.decode(answer_tokens).strip()
reward = 1.0 if answer == "4" else -1.0
prompt_targets = len(prompt_tokens) - 1
datum = types.Datum(
model_input=types.ModelInput.from_ints(prompt_tokens + answer_tokens[:-1]),
loss_fn_inputs={
"target_tokens": prompt_tokens[1:] + answer_tokens,
"logprobs": [0.0] * prompt_targets + logprobs,
"advantages": [0.0] * prompt_targets + [reward] * len(answer_tokens),
},
)
forward = training.forward_backward([datum], loss_fn="importance_sampling")
optimizer = training.optim_step(types.AdamParams(learning_rate=1e-5))
print(f"Answer: {answer!r}, reward: {reward}")
print("Training metrics:", forward.result(timeout=3600).metrics)
print("Optimizer metrics:", optimizer.result(timeout=3600).metrics)
updated_sampling = training.save_weights_and_get_sampling_client()
updated = (
updated_sampling.sample(
prompt=prompt,
num_samples=1,
sampling_params=params,
)
.result(timeout=3600)
.sequences[0]
)
print("After update:", tokenizer.decode(updated.tokens))For a dedicated full-parameter fine-tuning (FFT) run, use a scoped run, which brings up a single-tenant
server for the duration of the with block:
import spindle
import tinker
from spindle.engines import qwen3_5_4b_full_64k
engine = qwen3_5_4b_full_64k()
with spindle.run(engine=engine) as (url, api_key):
service = tinker.ServiceClient(base_url=url, api_key=api_key)
training = spindle.create_full_training_client(service, engine.model)
# Train and sample through the Tinker SDK here.See scoped runs for recovery and custom engines.
Since the server is just a Modal App, everything from the compute allocation, training backend details, and inference settings is fully customizable.
For example, this Qwen3.5-9B config demonstrates some settings you might change based on your training workload.
# config.py
from spindle.configuration import BaseConfig
class Config(BaseConfig):
model = "Qwen/Qwen3.5-9B"
name = "qwen35-9b-lora-16k"
max_context_length = 16_384
backend = "miles"
trainer_gpu = "H100"
trainer_gpus_per_node = 4
trainer_cpu = 16
trainer_memory_mib = 65_536
trainer_max_clients_per_instance = 6
miles_cfg = {
"model_type": "qwen3.5-9B",
"tensor_model_parallel_size": 4,
"max_lora_slots": 6,
"max_lora_rank": 32,
"default_lora_alpha": 32,
"target_modules": [
"linear_qkv",
"linear_proj",
"linear_fc1",
"linear_fc2",
"output_layer",
],
"max_tokens_per_gpu": 16_384,
"cli_options": {
"recompute_granularity": "full",
"recompute_method": "uniform",
"recompute_num_layers": 1,
},
}
inference_gpu = "H200"
inference_min_replicas = 2
inference_max_replicas = 8
sglang_cfg = {
"tp_size": 1,
"mem_fraction_static": 0.8,
"max_running_requests": 32,
"max_queued_requests": 8,
"max_loaded_loras": 64,
"max_loras_per_batch": 8,
}
config = Config()To deploy the above config, it's just:
spindle deploy config.pyAfter a training script exits, session heartbeats stop and Spindle's periodic cleaner reclaims idle training models and their latest sampler pools. Check that cleanup has finished in the Modal dashboard or with:
modal app listTo tear down the deployment, stop its spindle-fft-... sampler apps first, then the frontend app (spindle
by default, or whatever you passed to --app) with modal app stop <app-id>. Stopping the frontend does not
stop sampler apps.
Refer to the docs for design and for more advanced features when working with either the full-parameter or LoRA paths:
Read Working with Full Fine-Tunes for full training, or Working with Multi-LoRA for shared Miles adapters, batch submission, scheduling, and sampling. See scoped runs for recovery and custom engines on the full-parameter path.
See the raw Tinker RL example for sampling and a toy
policy update. Copy examples you want to run into your project; repository
scripts/ are not installed with the package.
The W&B RL example extends it to a multi-step
loop that logs reward, response length, and Spindle's training metrics to Weights
& Biases from the client side; tinker-cookbook users can instead set
wandb_project/wandb_name on the cookbook Config.
See Design for the control-plane, training-engine, and sampling architecture.
See Profiling for how to enable the torch.profiler trace of
a training step and read it in Perfetto.
See Observability for OTLP export to Datadog or a custom destination, experiment labels, and the complete span/metric inventory.
See FFT validation and LoRA validation for end-to-end training runs we've done with both parameterizations. The Codeforces codegolf example provides a larger-scale e2e code-RL training run, which trains Qwen3.5-9B with GRPO or TailRL advantages for correctness and short solutions. It includes a sandboxed judge, checkpoint recovery, and commands to continue a checkpoint with a different reward or advantage estimator, as well as pass@k and best-of-k evaluation.