vivago.ai

Unleash Creativity with Gen AI

A revolutionary AI-powered platform that brings professional-grade creative visual design within reach of everyone.

Sharper AI Videos Ahead: HiDream.ai’s GenVE Framework Accepted at ICCV

When it comes to AI-generated video, clarity and fidelity have always been tough nuts to crack. Now, a team from HiDream.ai has hit a major milestone: their GenVE (Generative Video Enhancement) framework has been officially accepted for presentation at ICCV 2025.

Research Background and Motivation

  1. Lack of Details: Current video enhancement technologies mostly rely on low-quality videos themselves, making it difficult to generate rich and realistic fine details and textures out of limited information. This causes AI-generated videos to often appear blurry or lack vividness in local areas, failing to meet users’ demands for high-quality visual content.
  2. Advantages of Images: Due to their static nature, high-quality images can usually capture and retain clearer and more abundant details than videos of the same resolution, and are not affected by degradation factors such as motion blur that are common in videos. GenVE cleverly leverages this advantage, using high-quality images as one of the “prior knowledge” sources to provide strong visual cues for video detail generation.
  3. Dual Alignment: The core of GenVE lies in achieving “global semantic alignment” and “local texture alignment.” Global semantic alignment ensures that the overall layout of the video is consistent with the reference image, avoiding content misalignment; local texture alignment focuses on pixel-level fine texture matching, ensuring that the video also has high-quality details at the microscopic level, thereby comprehensively enhancing the visual expressiveness of the video.

How GenVE Works

Semantic Alignment

When high-quality reference images of similar scenes are lacking, GenVE proposes to utilize advanced image diffusion models to upscale the key frames of the input video, generating a high-quality image as a semantic reference. This process employs a forward-backward diffusion mechanism combined with weighted summation to ensure that the generated image is highly consistent with the original video frames in terms of overall layout and high-level semantics, laying the foundation for subsequent detail enhancement.

Texture Alignment

GenVE integrates a video controller into the 3D-UNet of the video diffuser and introduces the Locally-Aware Cross-Attention (LCA) module. The LCA can accurately capture local texture details in high-quality image references and effectively transfer them to corresponding regions of low-quality videos. This mechanism enables the model to more precisely generate and repair fine textures in videos, such as hair and clothing wrinkles.

Enhancement Strategies

To improve the model’s robustness and generalization ability in complex scenarios, GenVE introduces multiple conditional enhancement strategies. The noise enhancement strategy is used to balance video quality and fidelity; the temporal enhancement strategy strengthens the temporal consistency of the video through optical flow information; the mask condition strategy prompts the model to learn and utilize different texture features in images more comprehensively by randomly masking parts of the image references, thereby enhancing its feature selection ability.

Experimental Results

Quantitative Leadership: On real-world video datasets such as YouHQ40 and VideoLQ, as well as the self-built AIGC-Vid synthetic video dataset, GenVE has achieved significant advantages across multiple mainstream video quality evaluation metrics including MUSIQ, DOVER, and CLIP-IQA. It comprehensively outperforms existing state-of-the-art video enhancement methods, fully demonstrating its exceptional performance and versatility.

Qualitative Excellence: Visual comparison charts clearly demonstrate GenVE’s strong capabilities. Whether in terms of the clarity of house windows or the naturalness of human faces, GenVE successfully synthesizes finer and more realistic details and textures. Even when dealing with extremely degraded videos, GenVE can accurately capture main structures and reasonably “envision” logical local content, making the videos look completely new.

Temporal Consistency: GenVE excels in temporal consistency. Through comparisons of temporal profile graphs, it is evident that videos generated by GenVE feature smoother and more seamless transitions between frames, effectively avoiding blurriness or discontinuities common in traditional methods. This is attributed to the motion-based conditional enhancement strategy introduced by GenVE, which ensures the coherence and natural fluency of the video.

Qualitative Results


The paper Aligning Global Semantics and Local Textures in Generative Video Enhancement, published at ICCV 2025, explores a new generative pathway for diffusion model-based video enhancement. By cleverly leveraging high-quality image references, GenVE achieves dual alignment of semantics and textures between input low-quality videos and generated high-quality videos, further alleviating the challenge of insufficient detail generation in videos.

In the future, GenVE will continue to advance video generation technology toward higher quality and more precise control, unlocking more possibilities for video creation.

Finally, we’d love for you to experience the future of AI video firsthand.w Welcome to try HiDream.ai’s flagship product: vivago!

Contacts

Company: HiDream.ai

Contact Person: Yuechong Zhai

Email: info@hidream.ai

Website: hidream.ai

Telephone: +86 13718564372

City: Beijing/Shanghai/Hefei

Discover more from vivago.ai

Subscribe now to keep reading and get access to the full archive.

Continue reading