CORTEXA
← Browse

Saurabh Agarwal

2 papers indexed

arxivcs.DCcs.LG2026-07-05

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

Nitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella

Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attention prevents exact autoregressive-style KV caching, since committing one position shifts the KV acti…

View free PDFSource page