Generating HDR Video from SDR Video
1Sony AI, Japan 2York University, Canada 3ISTI-CNR, Italy 4University of Toronto, Canada
SIGGRAPH Asia, 2026
The high dynamic range (HDR) video ecosystem is approaching maturity, but the problem of upconverting legacy standard dynamic range (SDR) videos persists without a convincing solution. We propose a framework for HDR video synthesis from casual SDR footage by leveraging a large-scale generative video model. We introduce a Multi-Exposure Video Model (MEVM) that can predict exposure-bracketed linear SDR video sequences from a single nonlinear SDR video input. We further propose a learnable Video Merging Model (VMM) that merges the predicted exposure-bracketed video into a high-quality HDR sequence while preserving detail in both shadows and highlights. Extensive experiments, quantitative and qualitative evaluation, and a user study demonstrate that our approach enables robust HDR conversion for in-the-wild examples from casual consumer videos and even cinematic footage. Finally, our model can support HDR synthesis pipelines built upon existing SDR generative video models.
Scroll down to view results.
Best experienced in Google Chrome
In-the-Wild Videos
Our method can even be applied to your own footage! Here we test our method on our own video albums and other arbitrarily sourced Internet video. Our method can generate clipped highlight regions and generate noisy quantized shadow regions.
Cinema Moments
We demonstrate our method on cinema footage. Our pipeline produces temporally coherent HDR outputs that enhance detail in both bright and dark regions without amplifying compression or grain noise.
Text to HDR Video Generation
Our pipeline can be chained with an off-the-shelf SDR video generation model to form a fully generative HDR video pipeline. We prompt a text-to-video model with exposure-related keywords such as “very bright”, “sunny day”, “very dark”, and “no lights” to produce an SDR input, which is then lifted to HDR; producing physically plausible dynamic-range expansion in both highlights and shadows while preserving temporal coherence.
Benchmarking our Method
Comparisons against baselines on the Stuttgart and UBC HDR benchmarks. Ground-truth HDR is available for each scene, enabling direct comparison. Switch the input exposure to inspect how each method handles over- and under-exposed inputs.
Limitations
Our fixed three-bracket exposure strategy cannot always recover the full dynamic range of high-contrast scenes, leaving some outputs with residual clipping after generation. We include representative failure cases to show where the approach breaks down and guide future improvements.