ReportGem ReportGem

Academic paper

Easper: An Accessible ASR Pipeline for Language Documentation

Authors: Aso Mahmudi, Ting Dang, Ekaterina Vylomova, Nick ThiebergerPublished: 2026-08-12Paper ID: 2608.11629Category: cs.CLLicense: CC BY 4.0

Abstract

Audio transcription is a critical bottleneck in language documentation. While multilingual Automatic Speech Recognition (ASR) models like Whisper offer solutions, field linguists often lack the expertise to utilise them. We present Easper, an open-source, no-code workflow enabling linguists to iteratively fine-tune ASR models via cloud resources directly from ELAN annotations. Deploying ASR also raises a cold start problem: deciding which recordings to transcribe first to bootstrap an accurate model. Using Easper, we evaluate transcription prioritisation strategies on three Vanuatu languages (Bislama, Nafsan, Nguna). We fine-tune models by recording session, comparing Character Error Rate trajectories when prioritising acoustic cleanliness versus linguistic richness. We demonstrate that prioritising lexically rich narratives and increasing acoustic-phonetic repetition, even in noisy environments, leads to faster improvements in transcription quality.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader