Loading comparison…
FlashVSR+
FlashVSR+
Efficient VSR. For everyone.
One-step restoration that makes every frame feel present.
01 / 4K video super-resolution
More detail.
More real.
FlashVSR+ turns low-resolution video into crisp, natural 4K detail in one step. Drag across the frame and see the difference in motion.
Consumer GPU. Full ambition.
4K restoration.
One RTX 4090.
Your local workstation.
01 / PROMPT
_
Open model · local workflow…
Bringing every frame into focus…
03 / In motion
Real World.
Natural detail. Everyday motion. Explore the difference, frame by frame.
Loading comparison…
04 / In motion
AIGC.
Imagined worlds, in finer detail. Explore super-resolution for generated video.
Loading comparison…
05 / Rediscover
Revival.
Bring old footage back to life. Restore familiar faces and moments, with the original sound preserved.
Preparing video…
Original audio Sound is off by default. Turn it on to relive the moment.
This is only the beginning.
FlashVSR+
Explore the research ↓
Abstract
Recent video super-resolution (VSR) models achieve efficient one-step generation through diffusion distillation. However, high-resolution inference remains costly, and distilled models are difficult to align with human preferences. These challenges call for improvements in both model architecture and training strategy. We present FlashVSR+ to address both limitations.
On the architecture side, we introduce hybrid attention based on the strong locality observed in pretrained VSR models, while retaining sparse global interaction for long-range consistency. It achieves up to 2× DiT speedup and 45% lower peak memory, enabling spatially untiled 4K VSR on a single 24GB RTX 4090. On the training side, we propose Reward-Yield Distribution Matching Distillation (Ry-DMD) for preference-aligned one-step generation. Ry-DMD uses reward-weighted teacher rollouts to incorporate black-box rewards into distribution matching, without backpropagating through the reward model. We further show that Ry-DMD generalizes beyond VSR to T2I generation. Together, FlashVSR+ delivers stronger restoration quality and substantially improves high-resolution efficiency, with consistent gains in user preference.
Three-stage training
Method Overview
A 14B full-attention teacher, a 1.3B hybrid-attention student, and reward-guided distillation form the FlashVSR+ pipeline.

Teacher
Train a 14B full-attention multi-step model with flow matching.
Hybrid student
Adapt a 1.3B backbone with local attention and sparse global blocks.
One-step model
Distill the student with distribution matching, pixel loss, and Ry-DMD.
Efficient high-resolution VSR
Hybrid Attention
Pretrained VSR attention is strongly concentrated around nearby space-time blocks. FlashVSR+ computes those interactions locally, while keeping periodic global blocks for distant context.

Key observations
Just ~1% of valid key blocks lie in the local neighborhood, yet they receive over 20× their proportional share of attention mass.
After hybrid training, the retained global blocks take on more global context, with attention mass spread more broadly across space and time.
Up to 2× faster DiT inference, with 45% lower peak GPU memory.
Reward-Yield Distribution Matching Distillation
Ry-DMD
Ry-DMD samples several teacher continuations from a noisy student state, scores their final outputs, and shifts the teacher target toward higher-reward samples. The reward only needs to return a scalar.
Core equations
Higher-reward teacher samples get more weight.
Compare reward-weighted and ordinary rollout means.
Inject the shift into DMD without differentiating through the reward.

Results

Efficiency comparison
Inference on 89-frame videos. The table reports GPU time and peak memory on an NVIDIA H200 equivalent GPU; the chart reports end-to-end wall-clock time on RTX 4090.
| Method | Params (B) | GPU time (s) ↓ | Peak memory (GB) ↓ | ||||
|---|---|---|---|---|---|---|---|
| 1080p | 2K | 4K | 1080p | 2K | 4K | ||
| DOVE | 5.0 | 89.38 | 165.3 | 374.7 | 26.54 | 26.22 | 26.56 |
| PS-SR | 1.3 | 228.4 | 1123.1 | 1751.4 | 47.79 | 50.63 | 50.63 |
| SeedVR2 | 3.0 | 51.46 | 157.8† | 410.5† | 111.21 | 49.57† | 51.76† |
| Ours-Base‡ | 1.3 | 7.91 | 17.28 | 46.34 | 13.11 | 21.29 | 45.00 |
| Ours-Turbo | 1.3 | 6.63 | 12.31 | 30.59 | 8.72 | 13.23 | 26.33 |

† SeedVR2 uses tiled inference at 2K and 4K because full-frame inference runs out of memory. ‡ Ours-Base shares FlashVSR's architecture and inference procedure on the H200 equivalent GPU; FlashVSR's attention operators are unsupported on RTX 4090.
Generalization beyond VSR
Ry-DMD for Image Generation
On 4-step SD3.5-M, Ry-DMD raises OCR accuracy from 0.501 to 0.824 at the same four-step sampling budget. Text becomes more legible while the scene remains intact.

Cite our paper
@article{zhuang2025flashvsr,
title={FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution},
author={Zhuang, Junhao and Guo, Shi and Cai, Xin and Li, Xiaohui and Liu, Yihao and Yuan, Chun and Xue, Tianfan},
journal={arXiv preprint arXiv:2510.12747},
year={2025}
}