ReportGem ReportGem

Academic paper

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Authors: Yejin Jeon, Marie Maltais, Virginia Ceccatelli, Min Ma, David Ifeoluwa AdelaniPublished: 2026-08-11Paper ID: 2608.10359Category: cs.SDLicense: CC BY 4.0

Abstract

As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader