Vladislav Bargatin*
·
Alexander Yakovenko*
·
Khaled Abud
·
Dmitriy Vatolin
*Equal contribution
Getting started · Checkpoints · Training
FreeFlow estimates optical flow using a hierarchical transformer with no task-specific inductive biases, and achieves state-of-the-art results on Sintel, KITTI, and Spring.
git clone https://github.com/msu-video-group/freeflow.git
cd freeflow
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e '.[examples,eval]'Requires Python ≥3.10 and PyTorch ≥2.5. See uv setup for environment details. Use a CUDA-enabled PyTorch installation for GPU inference.
Using Conda instead
conda create -n freeflow python=3.12 pip
conda activate freeflow
python -m pip install -e '.[examples,eval]'Run these commands from the cloned repository.
Optional CUDA RoPE extension
With a CUDA toolkit matching your PyTorch build and a C++ compiler:
uv pip install setuptools wheel ninja
uv pip install --no-build-isolation ./freeflow/models/curopeThe original CroCo code selects this extension when importable, otherwise PyTorch. For CPU inference, use the PyTorch fallback without the extension.
Weights are on Hugging Face: FreeFlow-S, FreeFlow-M, and FreeFlow-L.
For general-purpose optical flow, start with FreeFlow-L TaTSKH-HQ.
| Size | Depth | Width | Heads | Parameters |
|---|---|---|---|---|
| S | 4 | 256 | 4 | 34,575,109 |
| M | 6 | 384 | 6 | 102,456,965 |
| L | 8 | 512 | 8 | 230,721,029 |
The stages follow pretraining → TaTSKH → TaTSKH-HQ → benchmark fine-tuning. Sintel, KITTI, and Spring are separate fine-tunes.
Run the example with the included frames:
import numpy as np
import torch
from PIL import Image
from freeflow import FreeFlow
model = FreeFlow.from_pretrained("a-yakovenko/optical-flow-FreeFlow-L-TaTSKH-HQ").eval().cuda()
images = torch.stack([
torch.from_numpy(np.array(Image.open(path).convert("RGB"))).permute(2, 0, 1)
for path in ["assets/frame1.jpg", "assets/frame2.jpg"]
])[None]
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
flow = model(images) # [1, 2, H, W], float32
np.save("flow.npy", flow[0].permute(1, 2, 0).cpu().numpy())Inputs are RGB [B,2,3,H,W] tensors in [0,255]. The model estimates optical flow, returning a [B,2,H,W] array. Normalization and padding are handled automatically, and inference runs at native resolution.
For a flow array and color visualization:
freeflow-images assets/frame1.jpg assets/frame2.jpg \
--model a-yakovenko/optical-flow-FreeFlow-L-TaTSKH-HQ --output outputs/flowThe command supports --device cpu --precision fp32.
Place the datasets under flow_datasets/{sintel,kitti,spring} using their original layouts, then run:
freeflow-submit --dataset sintel --model_path checkpoints/S/sintel \
--config_dir freeflow/flow/configs --output_dir outputs/submissionsuv pip install -e '.[train]'
python -m freeflow.training.pretrain --help
python -m freeflow.flow.train --helpSee TRAINING.md for data configuration and checkpoint formats. Training code is included but has not yet been verified after the code release. Use FreeFlowPretraining.from_pretrained(...) to load a pretraining checkpoint for reconstruction.
@inproceedings{bargatin2026freeflow,
title={FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation},
author={Bargatin, Vladislav and Yakovenko, Alexander and Abud, Khaled and Vatolin, Dmitriy},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}Built on CroCo / CroCo v2 and MEMFOF, with the mixture-of-Laplace objective from SEA-RAFT. The implementation retains its development name, CroCo-Pro, in some internal modules. We thank the authors of these projects and the upstream components credited in LICENSE.
Code is distributed under CC BY-NC-SA 4.0.