blog
-
MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models
This blog post offers an introduction to MDM-Prime-v2, which scales the MDM-Prime framework to 1.1B-parameter models trained on 540B tokens. We introduce two zero-FLOP, lookup-table-based techniques, index shuffling and binary encoding, that resolve how sub-tokens should be formed and sized. This approach reaches 49.42% average accuracy on eight commonsense reasoning benchmarks, beating the MDM baseline (SMDM) by 4.87 points.
-
Dependency Breaks Validity of Loss Functions in Masked Diffusion Models
The MDM loss function is only a valid variational bound (i.e., an upper bound on negative log-likelihood) when the model factorizes token predictions independently. When dependencies between tokens are introduced into the parameterization — either through energy-based corrections (EDLM) or sub-token joint modeling (MDM-Prime) — the loss can fall below the data entropy, meaning it is no longer a valid bound and cannot be compared to perplexity or log-likelihood.
-
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
This blog post offers an introduction to MDM-Prime, a generalized masked diffusion model (MDM) that enables partially unmasked tokens during sampling. We begin with an review of MDMs and their limitations. Then, we explore a Partial masking scheme (Prime) that introduces intermediate token states between masked and unmasked representations. Finally, we present experimental results to demonstrate the effectiveness of MDM-Prime.
-
Maximum Entropy Reinforcement Learning via Energy-Based Normalizing Flow
This blog post offers an introduction to our proposed MEow algorithm. We begin with an review of MaxEnt RL and EBFlow. Then, we explore the connections between these models by introducing MEow. Finally, we present experimental results to demonstrate the effectiveness of the proposed method.
-
Training Energy-Based Normalizing Flow with Score-Matching Objectives
This blog post offers an introduction to our proposed EBFlow modeling method. First, we begin with an overview of flow-based and energy-based models. Then, we explore the connections between these models by introducing EBFlow. Next, we present experimental results to demonstrate the effectiveness of the proposed method. Finally, we discuss several implications of the EBFlow formula.