ReportGem ReportGem

Academic paper

LoopMTP: A looped transformer guided by latent multi-token prediction

Authors: Behzad Shomali, Markus Frey, David Berghaus, Joachim Koehler, Mehdi AliPublished: 2026-08-04Paper ID: 2608.03624Category: cs.CLLicense: CC BY 4.0

Abstract

Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across $T$ iterations, they attain the effective depth and reasoning capabilities of larger models at a fixed parameter count. Yet existing approaches suffer from latent overthinking and undifferentiated computation, largely because intermediate representations receive no guidance across loops. Multi-token prediction (MTP) supplies exactly the dense, forward-looking supervision the loop is missing. We propose \textsc{LoopMTP}, which links the two through a structural correspondence in latent space: a model that loops $T$ times can anticipate $T$ future tokens. \textsc{LoopMTP} realizes this by softly aligning the hidden state of loop $t$ with the embedding of the token $t$ steps ahead, while a lightweight gate preserves useful information across iterations. \textsc{LoopMTP} improves average accuracy by up to 8.1\% (relative) over the non-looped baseline, with training remaining stable for up to 15 loops.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader