AI Resonance · 阅读最新日报 · AI 入门推荐

2026-10-02 · America/Los_Angeles · 论文 · #19

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds…

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

热度 48.6 / 100;排名与评分保留该期记录。