Erdoğan, Hakan and Grais, Emad Mounir (2010) Semi-blind speech-music separation using sparsity and continuity priors. In: 20th International Conference on Pattern Recognition (ICPR 2010), Istanbul, Turkey
PDF
erdogan10sbs.pdf
Download (146kB)
erdogan10sbs.pdf
Download (146kB)
Official URL: http://dx.doi.org/10.1109/ICPR.2010.1129
Abstract
In this paper we propose an approach for the problem of single channel source separation of speech and music signals. Our approach is based on representing each source's power spectral density using dictionaries and nonlinearly projecting the mixture signal spectrum onto the combined span of the dictionary entries. We encourage sparsity and continuity of the dictionary coefficients using penalty terms (or log-priors) in an optimization framework. We propose to use a novel coordinate descent technique for optimization, which nicely handles nonnegativity constraints and nonquadratic penalty terms. We use an adaptive Wiener filter, and spectral subtraction to reconstruct both of the sources from the mixture data after corresponding power spectral densities (PSDs) are estimated for each source. Using conventional metrics, we measure the performance of the system on simulated mixtures of single person speech and piano music sources. The results indicate that the proposed method is a promising technique for low speech-to-music ratio conditions and that sparsity and continuity priors help improve the performance of the proposed system.
Item Type: | Papers in Conference Proceedings |
---|---|
Uncontrolled Keywords: | semi-blind signal separation , single channel speech-music separation , sparsity |
Subjects: | T Technology > TK Electrical engineering. Electronics Nuclear engineering Q Science > QA Mathematics > QA075 Electronic computers. Computer science |
Divisions: | Faculty of Engineering and Natural Sciences > Academic programs > Electronics Faculty of Engineering and Natural Sciences |
Depositing User: | Hakan Erdoğan |
Date Deposited: | 10 Dec 2010 15:00 |
Last Modified: | 26 Apr 2022 08:59 |
URI: | https://research.sabanciuniv.edu/id/eprint/15905 |