Towards More Flexible and Efficient Non-Autoregressive Language Models

A joint IESL talk on learnable insertion-order dynamics and relayed latent computation for non-autoregressive language models.

A joint talk on making non-autoregressive language models more flexible and efficient, covering two recent IESL group papers:

Discussion points:

  • How does learning the target generation order improve insertion-based diffusion without giving up tractable training?
  • What information is worth relaying across denoising steps, and how does truncated BPTT keep it trainable at scale?
  • Where do learnable-order and relayed-computation ideas complement each other for non-autoregressive language models?