From Data Scarcity to Data Care: Reimagining Language Technologies for Serbian and other Low-Resource Languages
arxiv.org·1d
🎙️Whisper
Preview
Report Post

View PDF

Abstract:Large language models are commonly trained on dominant languages like English, and their representation of low resource languages typically reflects cultural and linguistic biases present in the source language materials. Using the Serbian language as a case, this study examines the structural, historical, and sociotechnical factors shaping language technology development for low resource languages in the AI age. Drawing on semi structured interviews with ten scholars and practitioners, including linguists, digital humanists, and AI developers, it traces challenges rooted in historical destruction of Serbian textual heritage, intensified by contemporary issues that drive reductive, engineering first approaches prioritizing function…

Similar Posts

Loading similar posts...