vivago.ai

Unleash Creativity with Gen AI

A revolutionary AI-powered platform that brings professional-grade creative visual design within reach of everyone.

De-MAR: Fast & High-Quality Denoising Masked Autoregressive Generation Unveiled at ICCV2025

In recent years, autoregressive models, as the cornerstone of large language models (LLMs), have achieved remarkable success in the field of text generation. However, when applied to image generation, they often suffer from poor detail performance in generated images and much slower-than-expected inference speed due to their inherent one-way prediction mechanism and reliance on discretized tokens, making it difficult to meet users’dual demands for high quality and high efficiency simultaneously.

To tackle this challenge, the HiDream.ai team has proposed a brand-new framework called De-MAR, which enables denoising token prediction in masked autoregressive generative models (MAR) through a fundamental restructuring of the token prediction mechanism. This framework aims to address the core pain points of autoregressive models in visual generation, such as limited prediction context and inconsistency between training and inference, thus achieving breakthrough progress in both the quality and speed of generated images.

Research Background and Motivation

  1. Limited prediction context: When traditional masked autoregressive generative models predict masked image regions, their subsequent prediction modules (usually simple MLP structures) fail to fully utilize global contextual information, resulting in predicted details that lack logical consistency and realism.
  2. Training-inference discrepancy: Existing models use perfect “real” image patches as input during training, but during inference and generation, they rely on imperfect, error-prone image patches generated in their own previous steps. This inconsistency causes errors to accumulate continuously during the generation process, ultimately severely compromising the overall quality of the generated images.
  3. Dual token optimization: The core idea of De-MAR lies in the dual optimization of all tokens in the image. It not only predicts unknown “masked tokens” more accurately but also synchronously optimizes known “unmasked tokens”. By introducing a denoising mechanism, it gradually refines the “unmasked tokens”, thereby comprehensively improving the final performance of the generated images.

How De-MAR Works

  1. Diffusion Head Module: De-MAR introduces a Transformer-based Diffusion Head Module. Unlike the simple MLP heads in traditional models that only utilize local features, this Diffusion Head employs a cross-attention mechanism to receive and process conditional information provided by all tokens (including unmasked and masked positions) simultaneously. This global perspective enables more accurate prediction of masked tokens and enhances the Diffusion Head’s contextual awareness, thereby generating image content with logical coherence and rich details.
  2. Denoising Head Module: To address the issue of “training-inference discrepancy,” De-MAR has innovatively designed a Denoising Head Module. It injects noise into unmasked tokens during training to simulate the real scenario during inference. During the generation process, the Denoising Head is responsible for receiving these imperfect tokens. Through its internal token refinement branch and token evaluation branch, it dynamically denoises these tokens and assesses their quality, thereby improving the image quality of known regions and reducing error accumulation.

Experimental Results

  1. Leading quantitative metrics: On ImageNet and MS-COCO, two industry-recognized authoritative datasets, De-MAR has achieved significant advantages across multiple mainstream image quality assessment metrics. Particularly in terms of FID, a core metric for measuring the authenticity and diversity of generated images, De-MAR has reached top-level scores of 1.47 and 5.27 on the two datasets respectively, outperforming methods of similar autoregressive models and some classic diffusion models.
  2. Advantage in generation speed: While significantly improving image quality, De-MAR maintains extremely high generation efficiency. Benefiting from the inherent advantages of masked autoregressive generative models, the image generation inference time of De-MAR is 45% faster than that of widely recognized strong diffusion models (such as DiT-XL/2), breaking the inherent perception that high-quality generation is inevitably accompanied by high time costs.
  3. Excellent qualitative effects: The powerful generation capability of De-MAR is clearly evident in visual comparison charts. Compared with the baseline model MAR, images generated by De-MAR (such as animal fur and fruit textures) feature finer, more realistic details, fewer artifacts, and an overall more natural appearance.

The paper Denoising Token Prediction in Masked Autoregressive Models, published at ICCV 2025, has explored a new path that balances quality and efficiency for autoregressive image generation technology. Through its innovative denoising token prediction mechanism, De-MAR has subtly addressed two major challenges in existing models: insufficient utilization of context and training-inference discrepancy, further unleashing the great potential of autoregressive models in the field of visual generation.

In the future, the HiDream.ai team will continue to explore next-generation multimodal generative model architectures, committing to achieving multimodal content generation technology with higher quality, faster speed, and stronger controllability.

Finally, we’d love for you to experience the future of AI video firsthand. Welcome to try HiDream.ai’s flagship product: vivago!

Contacts
Company: HiDream.ai
Contact Person: Yuechong Zhai
Email: info@hidream.ai
Website: hidream.ai
Telephone: +86 13718564372
City: Beijing/Shanghai/Hefei

Discover more from vivago.ai

Subscribe now to keep reading and get access to the full archive.

Continue reading