This paper introduces audio-guided video scene editing, a new task that dynamically modifies the background of a video based on audio input while preserving the integrity of the foreground. To address this challenge, we propose AudioScenic, a specialized framework designed to achieve four core objectives: integrating audio semantics into video scenes, maintaining the foreground’s visual stability, ensuring alignment between audio condition and visual changes, and upholding temporal consistency throughout the video. At the heart of AudioScenic is a temporal-aware audio semantic injection mechanism, which effectively integrates audio conditions into the video background. To prevent undesired distortions in the foreground, we introduce SceneMasker, which ensures that foreground elements remain unaffected during editing. Furthermore, AudioScenic incorporates two key components: the Magnitude Modulator, which strengthens synchronization between auditory and visual elements over time, and the Frequency Fuser, which exploits the inherent frequency correlations between audio and video to enhance temporal coherence. With these innovations, AudioScenic generates visually diverse, seamlessly synchronized video scenes that adapt to the given audio while maintaining smooth frame transitions. Additionally, we introduce a new evaluation metric, the Semantic-Temporal Consistency Score (STC-Score), which offers a more comprehensive measurement of temporal consistency in video scene editing. Experimental results demonstrate that AudioScenic significantly outperforms existing methods, establishing it as a robust solution for synchronized video scene editing.