Nemotron-Labs-TwoTower

NVIDIA Research's open-weight diffusion language model, adapted from a frozen Nemotron-3-Nano-30B-A3B backbone, uses a two-tower design where one tower holds context and the other writes tokens in parallel, achieving ~2.4× throughput without retraining.