Skip to main content
Have a personal or library account? Click to login
Real Time Recognition Of Speakers From Internet Audio Stream Cover

Real Time Recognition Of Speakers From Internet Audio Stream

Open Access
|Sep 2015

References

  1. [1] S. Araki, T. Hori, M. Fujimoto, S. Watanabe, T. Yoshioka, T. Nakatani, and A. Nakamura. Online meeting recognizer with multichannel speaker diarization. In, pages 1697–1701, Nov 2010.
  2. [2] D. Blatt and A. Hero. On tests for global maximum of the log-likelihood function., 53(7):2510–2525, July 2007.
  3. [3] M. Bosi, K. Brandenburg, S. Quackenbush, L. Fielder, K. Akagiri, H. Fuchs, and M. Dietz. ISO/IEC MPEG-2 Advanced Audio Coding., 45(10):789–814, 1997.
  4. [4] M. Brookes. VOICEBOX: Speech Processing Toolbox for MATLAB, 2005.
  5. [5] J. Dattorro.. Lulu. com, 2008.
  6. [6] J. R. Hershey and R. A. Olsen. Approximating the Kullback Leibler divergence between gaussian mixture models. In, pages 317–320, 2007.
  7. [7] T. Jiang and J. Han. Map-based audio coding compensation for speaker recognition., 2:165, 2011.
  8. [8] R. D. Maesschalck, D. Jouan-Rimbaud, and D. Massart. The Mahalanobis distance., 50(1):1 – 18, 2000.
  9. [9] T. Marciniak, R. Weychan, A. Dabrowski, and A. Krzykowska. Speaker recognition based on short Polish sequences., pages 95–98, 2010.
  10. [10] T. Marciniak, R. Weychan, A. Dabrowski, and A. Krzykowska. Influence of silence removal on speaker recognition based on short Polish sequences., pages 159–163, 2011.
  11. [11] T. Marciniak, R. Weychan, A. Stankiewicz, and A. Dabrowski. Biometric speech signal processing in a system with digital signal processor., Vol. 62, nr 3:589–594, 2014.
  12. [12] S. Molau, M. Pitz, R. Schluter, and H. Ney. Computing Mel-frequency cepstral coefficients on the power spectrum. In, volume 1, pages 73–76, 2001.
  13. [13] K. Park, J.-S. Park, and Y.-H. Oh. GMM adaptation based online speaker segmentation for spoken document retrieval., 56(2):1123–1129, 2010.
  14. [14] Z. Piotrowski, J. Wojtun, and K. Kaminski. Subscriber authentication using GMM and tms320c6713dsp., (12a/2012):127–130, 2012.
  15. [15] A. Plinge and G. A. Fink. Online multi-speaker tracking using multiple microphone arrays informed by auditory scene analysis. In, pages 1–5, Sept 2013.
  16. [16] D. Reynolds. Gaussian mixture models., pages 659–663, 2009.
  17. [17] J. B. Tenenbaum, V. D. Silva, and J. C. Langford. A global geometric framework for nonlinear dimensionality reduction., 290(5500):2319–2323, 2000.
  18. [18] G. Wen, L. Jiang, and J. Wen. Using locally estimated geodesic distance to optimize neighborhood graph for isometric data embedding., 41(7):2226 – 2236, 2008.
  19. [19] R. Weychan, T. Marciniak, and A. Dabrowski. Analysis of differences between MFCC after multiple GSM transcodings., pages 24–29, 2012.
  20. [20] R. Weychan, T. Marciniak, A. Stankiewicz, and A. Dabrowski. Real time speaker recognition from internet radio., pages 128–132, 2014.
  21. [21] R. Weychan, A. Stankiewicz, T. Marciniak, and A. Dabrowski. Improving of speaker identification from mobile telephone calls. In, volume 429 of, pages 254–264. 2014.
DOI: https://doi.org/10.1515/fcds-2015-0014 | Journal eISSN: 2300-3405 (formerly 0867-6356) | Journal ISSN: 0867-6356
Language: English
Page range: 223 - 233
Published on: Sep 30, 2015
Published by: Poznan University of Technology
In partnership with: Paradigm Publishing Services

© 2015 Radoslaw Weychan, Tomasz Marciniak, Agnieszka Stankiewicz, Adam Dabrowski, published by Poznan University of Technology
This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 3.0 License.