SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
arxiv.org·1h
👁️Perceptual Coding
Preview
Report Post

Title:SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality

View PDF HTML (experimental)

Abstract:Objective speech quality assessment is central to telephony, VoIP, and streaming systems, where large volumes of degraded audio must be monitored and optimized at scale. Classical metrics such as PESQ and POLQA approximate human mean opinion scores (MOS) but require carefully controlled conditions and expensive listening tests, while learning-based models such as NISQA regress MOS and multiple perceptual dimensions from waveforms or spectrograms, achieving high correlation with subjective ratings yet remaining rigid: they do not support interactive, natural-language queries and do not natively …

Similar Posts

Loading similar posts...