VIDEO SUPER-RESOLUTION

FlashVSR+

FlashVSR+

Efficient VSR. For everyone.

One-step restoration that makes every frame feel present.

01 / 4K video super-resolution

More detail.
More real.

FlashVSR+ turns low-resolution video into crisp, natural 4K detail in one step. Drag across the frame and see the difference in motion.

960p input 4K output

Loading comparison…

Low resolutionFlashVSR+
LR
SRDrag to compare

Consumer GPU. Full ambition.

4K restoration.
One RTX 4090.

ThroughputRTX 4090 · 4K · FPS
0.18
0.50
0.98
5.5×faster than SeedVR2
GENERATE & ENHANCE LOCALLY

Your local workstation.

01 / PROMPT

01 YOUR PROMPT

_

02 / VIDEO GENERATIONMiniMax H3

Open model · local workflow…

LOW RESOLUTION OUTPUT
MINIMAX H3 · 1344 × 768
03 / VIDEO SUPER-RESOLUTIONFlashVSR+

Bringing every frame into focus…

FINAL HIGH RESOLUTION VIDEO
FLASHVSR+ · 3840 × 2160
FOLLOW THE WORKFLOW

03 / In motion

Real World.

Natural detail. Everyday motion. Explore the difference, frame by frame.

Loading comparison…

Low resolutionFlashVSR+
LR
SRDrag to compare

Demo 01 / 05

04 / In motion

AIGC.

Imagined worlds, in finer detail. Explore super-resolution for generated video.

Loading comparison…

Low resolutionFlashVSR+
LR
SRDrag to compare

Demo 01 / 06

05 / Rediscover

Revival.

Bring old footage back to life. Restore familiar faces and moments, with the original sound preserved.

Preparing video…

OriginalRestored
LR
SRDrag to compare

Original audio Sound is off by default. Turn it on to relive the moment.

This is only the beginning.

FlashVSR+

Explore the research ↓
FlashVSR+ paper teaser comparing restored video detail against baselines, with 4K efficiency results
FlashVSR+ restores fine detail and natural textures while improving 4K throughput.

Abstract

Recent video super-resolution (VSR) models achieve efficient one-step generation through diffusion distillation. However, high-resolution inference remains costly, and distilled models are difficult to align with human preferences. These challenges call for improvements in both model architecture and training strategy. We present FlashVSR+ to address both limitations.

On the architecture side, we introduce hybrid attention based on the strong locality observed in pretrained VSR models, while retaining sparse global interaction for long-range consistency. It achieves up to 2× DiT speedup and 45% lower peak memory, enabling spatially untiled 4K VSR on a single 24GB RTX 4090. On the training side, we propose Reward-Yield Distribution Matching Distillation (Ry-DMD) for preference-aligned one-step generation. Ry-DMD uses reward-weighted teacher rollouts to incorporate black-box rewards into distribution matching, without backpropagating through the reward model. We further show that Ry-DMD generalizes beyond VSR to T2I generation. Together, FlashVSR+ delivers stronger restoration quality and substantially improves high-resolution efficiency, with consistent gains in user preference.

Three-stage training

Method Overview

A 14B full-attention teacher, a 1.3B hybrid-attention student, and reward-guided distillation form the FlashVSR+ pipeline.

FlashVSR+ three-stage training pipeline: full-attention teacher pretraining, hybrid-attention initialization, and Ry-DMD distillation
Figure 1. FlashVSR+ training pipeline. Stage 3 uses reward feedback to modify the teacher term without reward gradients.
01

Teacher

Train a 14B full-attention multi-step model with flow matching.

02

Hybrid student

Adapt a 1.3B backbone with local attention and sparse global blocks.

03

One-step model

Distill the student with distribution matching, pixel loss, and Ry-DMD.

Efficient high-resolution VSR

Hybrid Attention

Pretrained VSR attention is strongly concentrated around nearby space-time blocks. FlashVSR+ computes those interactions locally, while keeping periodic global blocks for distant context.

Attention locality map and local attention mass across DiT blocks before and after hybrid training
Attention analysis. Compact local neighborhoods receive a disproportionate share of attention; retained global blocks become less local after training.

Key observations

Primarily local

Just ~1% of valid key blocks lie in the local neighborhood, yet they receive over 20× their proportional share of attention mass.

Periodically global

After hybrid training, the retained global blocks take on more global context, with attention mass spread more broadly across space and time.

2× DiT speedup on 4K

Up to 2× faster DiT inference, with 45% lower peak GPU memory.

Reward-Yield Distribution Matching Distillation

Ry-DMD

Ry-DMD samples several teacher continuations from a noisy student state, scores their final outputs, and shifts the teacher target toward higher-reward samples. The reward only needs to return a scalar.

Core equations

01 / Reward-tilted target
p0⋆(x0∣c)∝pT,0(x0∣c)exp⁡ ⁣(βr(x0,c))p_0^\star(x_0\mid c)\propto p_{\mathrm T,0}(x_0\mid c)\exp\!\bigl(\beta r(x_0,c)\bigr)

Higher-reward teacher samples get more weight.

02 / Local target shift
δm=∑j=1Jπjx0(j)−1J∑j=1Jx0(j)\delta m=\sum_{j=1}^{J}\pi_j x_0^{(j)}-\frac{1}{J}\sum_{j=1}^{J}x_0^{(j)}

Compare reward-weighted and ordinary rollout means.

03 / Teacher correction
sTRY=sT+λRYδmmax⁡(σ,εσ)s_{\mathrm T}^{\mathrm{RY}}=s_{\mathrm T}+\lambda_{\mathrm{RY}}\frac{\delta m}{\max(\sigma,\varepsilon_\sigma)}

Inject the shift into DMD without differentiating through the reward.

Conceptual Ry-DMD diagram showing a teacher distribution shifted toward higher-reward samples
Ry-DMD concept: reward-weighted teacher rollouts shift the target distribution toward desirable outputs.

Results

Qualitative comparison of FlashVSR+ against representative VSR methods on text, parrot texture, and animated facial detail
Qualitative comparison with representative VSR methods. Zoomed regions highlight text, fine textures, and facial details.

Efficiency comparison

Inference on 89-frame videos. The table reports GPU time and peak memory on an NVIDIA H200 equivalent GPU; the chart reports end-to-end wall-clock time on RTX 4090.

(a) NVIDIA H200 equivalent GPU
MethodParams (B)GPU time (s) ↓Peak memory (GB) ↓
1080p2K4K1080p2K4K
DOVE5.089.38165.3374.726.5426.2226.56
PS-SR1.3228.41123.11751.447.7950.6350.63
SeedVR23.051.46157.8†410.5†111.2149.57†51.76†
Ours-Base‡1.37.9117.2846.3413.1121.2945.00
Ours-Turbo1.36.6312.3130.598.7213.2326.33
RTX 4090 end-to-end wall-clock efficiency comparison across 1080p, 2K and 4K resolutions
(b) NVIDIA RTX 4090

† SeedVR2 uses tiled inference at 2K and 4K because full-frame inference runs out of memory. ‡ Ours-Base shares FlashVSR's architecture and inference procedure on the H200 equivalent GPU; FlashVSR's attention operators are unsupported on RTX 4090.

Generalization beyond VSR

Ry-DMD for Image Generation

On 4-step SD3.5-M, Ry-DMD raises OCR accuracy from 0.501 to 0.824 at the same four-step sampling budget. Text becomes more legible while the scene remains intact.

Comparison of baseline and Ry-DMD four-step SD3.5-M generations: candle, magnifying glass, and elevator sign with clearer text after Ry-DMD
Qualitative comparison on 4-step SD3.5-M. Top: baseline. Bottom: Ry-DMD. Prompts are shown in the figure.

Cite our paper

@article{zhuang2025flashvsr,
  title={FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution},
  author={Zhuang, Junhao and Guo, Shi and Cai, Xin and Li, Xiaohui and Liu, Yihao and Yuan, Chun and Xue, Tianfan},
  journal={arXiv preprint arXiv:2510.12747},
  year={2025}
}