Two Halves are More than One:
Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
Two complementary halves for high-quality image generation at roughly one full-backbone forward pass of compute.
π© Accepted by NeurIPS 2026
Zhen Guo Β· Rongyuan Wu Β· Qiaosi Yi Β· Chenxi Xie Β· Xinyu Wei Β· Lei Zhang
The Hong Kong Polytechnic University Β· OPPO Research Institute
News Β· Highlights Β· Overview Β· Results Β· Gallery Β· Preparation
Training Β· Inference Β· Contact Β· Citation Β· License Β· Acknowledgements
If you find this repository helpful, please kindly give it a star β.
- π Code released! Training and inference code is available for C2I and all three T2I backbones.
- π€ Checkpoints released! Download the C2I, SD3.5 Medium, FLUX.1-dev and Qwen-Image models, including the Unsplash variants.
PVD distills a teacher into two complementary, half-depth experts: the early expert establishes global structure, and the late expert refines detail. The experts run sequentially, once each, sharing the computation budget of approximately one full teacher-backbone forward pass.
- Near-teacher ImageNet quality at 1/500 of the sampling FLOPs. On ImageNet 256Γ256, PVD achieves FID-50K 1.48 Β· IS 295.89, approaching the multi-step LightningDiT teacher's FID of 1.35 with normalized sampling FLOPs of 1.00 vs. 500.00.
- Strong text-to-image results across three backbones. PVD reaches GenEval 0.6948 / 0.6532 / 0.8846 on SD3.5 Medium / FLUX.1-dev / Qwen-Image, respectively, under the same approximate single-forward compute budget.
- Smaller experts, lower memory use. Across the T2I backbones, PVD uses 49.10β50.89% fewer active parameters and 45.76β48.36% less peak VRAM than the multi-step teachers.
We first extract and align a half-depth backbone from the teacher. Each phase expert is then distilled on a shorter time interval with a local mean-velocity objective. For T2I, phase-specific adversarial supervision further improves perceptual detail. At inference, the late expert consumes the early expert's intermediate state.
FID is computed on 50K class-conditional samples. Lower FID and higher Inception Score (IS) are better. N_flops is sampling FLOPs normalized by one full teacher-backbone forward pass; the multi-step LightningDiT teacher is shown for context.
| Method | Params | N_flops |
FID-50K β | IS β |
|---|---|---|---|---|
| LightningDiT (multi-step) | 675M | 500.00 | 1.35 | 295.30 |
| MeanFlow | 676M | 1.00 | 3.43 | β |
| Ξ±-Flow | 675M | 1.00 | 2.58 | β |
| FACM | 675M | 1.00 | 1.76 | β |
| iMF | 610M | 1.00 | 1.72 | 282.00 |
| PVD | 676M | 1.00 | 1.48 | 295.89 |
PVD's FID is only 0.13 above the multi-step teacher's, with comparable IS, at 1/500 of its normalized sampling FLOPs.
Each table compares a multi-step teacher, a representative accelerated baseline and PVD. Higher is better for all quality metrics.
| Method | GenEval β | DPG β | WISE β | Qwen-Image-Bench β | Aesthetic β | PickScore β | ImageReward β |
|---|---|---|---|---|---|---|---|
| Teacher | 0.6829 | 84.49 | 0.45 | 43.19 | 5.7212 | 21.45 | 0.6127 |
| SWD | 0.6286 | 79.23 | 0.36 | 31.44 | 4.9440 | 20.53 | 0.3578 |
| PVD | 0.6948 | 81.81 | 0.40 | 35.51 | 5.4618 | 20.84 | 0.4947 |
| Method | GenEval β | DPG β | WISE β | Qwen-Image-Bench β | Aesthetic β | PickScore β | ImageReward β |
|---|---|---|---|---|---|---|---|
| Teacher | 0.6675 | 83.95 | 0.50 | 43.52 | 5.9286 | 21.81 | 0.8139 |
| SWD | 0.6285 | 82.43 | 0.37 | 39.71 | 5.3846 | 21.22 | 0.6562 |
| PVD | 0.6532 | 82.83 | 0.44 | 43.14 | 5.7145 | 21.32 | 0.6747 |
| Method | GenEval β | DPG β | WISE β | Qwen-Image-Bench β | Aesthetic β | PickScore β | ImageReward β |
|---|---|---|---|---|---|---|---|
| Teacher | 0.8720 | 89.10 | 0.64 | 48.78 | 5.8584 | 21.90 | 0.9852 |
| TwinFlow | 0.8338 | 86.75 | 0.57 | 46.16 | 5.5880 | 21.42 | 0.9042 |
| PVD | 0.8846 | 86.54 | 0.59 | 46.85 | 5.7526 | 21.66 | 0.9063 |
The cost figures below are from the paper's evaluation; active parameters and peak VRAM are shown as teacher β PVD, with the reduction relative to the teacher.
| Backbone | PVD N_flops |
Active params (B) | Peak VRAM (GiB) |
|---|---|---|---|
| SD3.5 Medium | 0.99 | 2.24 β 1.10 β50.89% |
4.99 β 2.67 β46.49% |
| FLUX.1-dev | 1.01 | 11.90 β 5.98 β49.75% |
22.64 β 12.28 β45.76% |
| Qwen-Image | 1.02 | 20.43 β 10.40 β49.10% |
38.42 β 19.84 β48.36% |
ImageNet Β· Text-to-image Β· Unsplash
PVD samples for 256Γ256 class-conditional generation:
PVD samples for 1024Γ1024 text-to-image, grouped by backbone:
Further training on Unsplash improves lighting, texture and realism, showing that PVD can effectively adapt the visual characteristics of high-quality training data into the generation process beyond teacher imitation.
Use a Python environment with a CUDA-enabled GPU. Clone the repository and install its dependencies:
git clone https://github.com/PolyU-VCLab/PVD.git
cd PVD
pip install -r requirements.txtRun subsequent commands from the PVD repository root.
All seven student variants are available on Hugging Face. Each model link opens its weight folder. Place weights/ and cache/ in this repository root, preserving the directory names below.
PVD/
βββ weights/ # PVD student checkpoints and adapters
βββ cache/ # C2I decoder, statistics and teacher checkpoint
βββ models/ # T2I base model components
βββ data/ # ImageNet latents and imageβcaption manifests
βββ outputs/ # Generated images and training runs
| Model / download | Local weight directory | Released components |
|---|---|---|
| C2I | weights/pvd_c2i/ |
model.safetensors |
| SD3.5 Medium | weights/pvd_sd35m/ |
part1.safetensors, part2.safetensors, model_config.json |
| SD3.5 Medium (Unsplash) | weights/pvd_sd35m_unsplash/ |
part1.safetensors, part2.safetensors, model_config.json |
| FLUX.1-dev | weights/pvd_flux/ |
backbone.safetensors, part1/ & part2/ LoRA adapters |
| FLUX.1-dev (Unsplash) | weights/pvd_flux_unsplash/ |
part1/ & part2/ adapters; shared FLUX backbone below |
| Qwen-Image | weights/pvd_qwenimage/ |
backbone.safetensors, part1/ & part2/ LoRA adapters |
| Qwen-Image (Unsplash) | weights/pvd_qwen_unsplash/ |
part1/ & part2/ adapters; shared Qwen backbone below |
The Unsplash FLUX and Qwen-Image variants reuse weights/pvd_flux/backbone.safetensors and weights/pvd_qwenimage/backbone.safetensors, respectively. Each LoRA folder contains adapter_model.safetensors and adapter_config.json. For C2I inference, download the decoder and latent statistics as cache/vavae-imagenet256-f16d32-dinov2.pt and cache/latents_stats.pt.
For T2I, place the base model's VAE, tokenizers and text encoders in models/stable-diffusion-3.5-medium/, models/FLUX.1-dev/, or models/Qwen-Image/, and pass that directory with --model-root.
Dataset sources and training assets are listed below.
| Stage | Data used in the paper | Source / local preparation |
|---|---|---|
| C2I training | ImageNet-256; FACM weights & statistics | HuggingFace |
| T2I backbone initialization | BLIP3o long captions | BLIP3o-Pretrain-Long-Caption |
| T2I phase distillation | BLIP3o-60k; also Echo-4o for Qwen-Image | BLIP3o-60k Β· Echo-4o-Image |
| Further T2I training | Unsplash photographs | Unsplash Dataset |
Download the FACM weights and statistics linked above. Place fid-50k-256.npz, latents_stats.pt and vavae-imagenet256-f16d32-dinov2.pt in cache/. Prepare ImageNet latents following Lightning-DiT.
For T2I, create a JSONL manifest with one record per image. image_path must be relative to IMAGE_ROOT; caption must contain the training text:
{"image_path": "images/000001.jpg", "caption": "A red apple on a wooden table"}Set DATA_FILE to this manifest and IMAGE_ROOT to its image root in the training commands. Text is encoded online by default.
Complete Preparation, then adjust the GPU count and local paths in the commands below to your setup.
| Task | Entry point | Loss implementation |
|---|---|---|
| C2I | train.py | losses/c2i.py |
| SD3.5 Medium | train_t2i_sd3.py | losses/sd35.py |
| FLUX.1-dev | train_t2i_flux.py | losses/flux.py |
| Qwen-Image | train_t2i_qwenimage.py | losses/qwenimage.py |
GPUS=8
export C2I_TEACHER_CHECKPOINT=cache/800ep-stg1.pt
accelerate launch --num_processes "$GPUS" --mixed_precision bf16 train.py \
--data-dir data/imagenet_latents \
--results-dir outputs/c2i \
--distill --intervals 0.4_0.6 --overlap 0.0Use scripts/train_t2i.sh with the full teacher transformer in each base-model directory. Train the two phases separately: run with PART=1, then repeat with PART=2. Accelerate / DeepSpeed configurations are in scripts/config/.
TASK=flux PART=1 GPUS=8 DATA_FILE=data/blip3o/train.jsonl IMAGE_ROOT=data/blip3o/images \
FLUX_MODEL_ROOT=models/FLUX.1-dev bash scripts/train_t2i.sh
TASK=sd35 PART=1 GPUS=8 DATA_FILE=data/blip3o/train.jsonl IMAGE_ROOT=data/blip3o/images \
SD35_MODEL_ROOT=models/stable-diffusion-3.5-medium bash scripts/train_t2i.sh
TASK=qwenimage PART=1 GPUS=8 DATA_FILE=data/blip3o/train.jsonl IMAGE_ROOT=data/blip3o/images \
QWENIMAGE_MODEL_ROOT=models/Qwen-Image bash scripts/train_t2i.shSingle-GPU inference with the prepared models. Outputs are saved to the path specified by --output.
# ImageNet class ID 207
python infer.py --task c2i --class-id 207 --output outputs/c2i.pngpython infer.py --task sd35 --model-root models/stable-diffusion-3.5-medium \
--prompt "A red apple on a wooden table" --output outputs/sd35.png
python infer.py --task flux --model-root models/FLUX.1-dev \
--prompt "A red apple on a wooden table" --output outputs/flux.png
python infer.py --task qwenimage --model-root models/Qwen-Image \
--prompt "A red apple on a wooden table" --output outputs/qwenimage.pngpython infer.py --task sd35_unsplash --model-root models/stable-diffusion-3.5-medium \
--prompt "A red apple on a wooden table" --output outputs/sd35_unsplash.png
python infer.py --task flux_unsplash --model-root models/FLUX.1-dev \
--prompt "A red apple on a wooden table" --output outputs/flux_unsplash.png
python infer.py --task qwenimage_unsplash --model-root models/Qwen-Image \
--prompt "A red apple on a wooden table" --output outputs/qwenimage_unsplash.png| Argument | Purpose |
|---|---|
--task |
Select a student variant from the model table |
--weights-dir |
Parent directory of the student weight folders; default: weights/ |
--model-root |
Base model directory; required for T2I |
--class-id |
ImageNet class ID for C2I (0β999) |
--part1_steps, --part2_steps |
T2I steps per phase; both default to 1 |
If you have any questions or suggestions, please feel free to contact: zhen-gz.guo@connect.polyu.hk.
If you find PVD useful, please consider citing our work:
@inproceedings{guo2026pvd,
title = {Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation},
author = {Guo, Zhen and Wu, Rongyuan and Yi, Qiaosi and Xie, Chenxi and Wei, Xinyu and Zhang, Lei},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}Code is licensed under the Apache License 2.0. Third-party code and components retain their respective licenses and copyright notices. Model weights are subject to the applicable upstream model licenses.
This repository is built upon FACM, Transformers, and Diffusers. We thank the authors for their awesome work!









