PAVAS: Physics-Aware Video-to-Audio Synthesis
arxiv.org·1h
🎧Vorbis Encoding
Preview
Report Post

Title:PAVAS: Physics-Aware Video-to-Audio Synthesis

View PDF

Abstract:Recent advances in Video-to-Audio (V2A) generation have achieved impressive perceptual quality and temporal synchronization, yet most models remain appearance-driven, capturing visual-acoustic correlations without considering the physical factors that shape real-world sounds. We present Physics-Aware Video-to-Audio Synthesis (PAVAS), a method that incorporates physical reasoning into a latent diffusion-based V2A generation through the Physics-Driven Audio Adapter (Phy-Adapter). The adapter receives object-level physical parameters estimated by the Physical Parameter Estimator (PPE), which uses a Vision-Language Model (VLM) to infer the moving-object mass and a segmen…

Similar Posts

Loading similar posts...