Refine Your Search

Search Results

Viewing 1 to 2 of 2
Technical Paper

A Unified Frequency Understanding of Image Corruptions and its Application to Autonomous Driving

2023-04-11
2023-01-0060
Image corruptions due to noise, blur, contrast change, etc., could lead to a significant performance decline of Deep Neural Networks (DNN), which poses a potential threat to DNN-based autonomous vehicles. Previous works attempted to explain corruption from a Fourier perspective. By comparing the absolute Fourier spectrum difference between corrupted images and clean images in the RGB color space, they regard the noise from some corruptions (Gaussian noise, defocus blur, etc.) as concentrating on the high-frequency components while others (contrast, fog, etc.) concentrate on the low-frequency components. In this work, we present a new perspective that unifies corruptions as noise from high frequency and thus propose an image augmentation algorithm to achieve a more robust performance against common corruptions. First, we notice the 1/fα statistical rule of the natural image's spectrum and the channels-wise differential sensitivity on the YCbCr color space of the Human Visual System.
Technical Paper

A Target-Speech-Feature-Aware Module for U-Net Based Speech Enhancement

2024-04-09
2024-01-2021
Speech enhancement can extract clean speech from noise interference, enhancing its perceptual quality and intelligibility. This technology has significant applications in in-car intelligent voice interaction. However, the complex noise environment inside the vehicle, especially the human voice interference is very prominent, which brings great challenges to the vehicle speech interaction system. In this paper, we propose a speech enhancement method based on target speech features, which can better extract clean speech and improve the perceptual quality and intelligibility of enhanced speech in the environment of human noise interference. To this end, we propose a design method for the middle layer of the U-Net architecture based on Long Short-Term Memory (LSTM), which can automatically extract the target speech features that are highly distinguishable from the noise signal and human voice interference features in noisy speech, and realize the targeted extraction of clean speech.
X