Leveraging native language speech for accent identification using deep Siamese networks

Siddhant, A and Jyothi, P and Ganapathy, S (2018) Leveraging native language speech for accent identification using deep Siamese networks. In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2017, 16 - 20 December 2017, Okinawa, pp. 621-628.

Preview

PDF
IEEE_ASRU 2017_2018_621-628_2018.pdf - Published Version
Download (341kB) | Preview

Official URL: https://doi.org/10.1109/ASRU.2017.8268994

Abstract

The problem of automatic accent identification is important for several applications like speaker profiling and recognition as well as for improving speech recognition systems. The accented nature of speech can be primarily attributed to the influence of the speaker's native language on the given speech recording. In this paper, we propose a novel accent identification system whose training exploits speech in native languages along with the accented speech. Specifically, we develop a deep Siamese network based model which learns the association between accented speech recordings and the native language speech recordings. The Siamese networks are trained with i-vector features extracted from the speech recordings using either an unsupervised Gaussian mixture model (GMM) or a supervised deep neural network (DNN) model. We perform several accent identification experiments using the CSLU Foreign Accented English (FAE) corpus. In these experiments, our proposed approach using deep Siamese networks yield significant relative performance improvements of 15.4 on a 10-class accent identification task, over a baseline DNN-based classification system that uses GMM i-vectors. Furthermore, we present a detailed error analysis of the proposed accent identification system.

Item Type:	Conference Paper
Publication:	2017 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2017 - Proceedings
Publisher:	Institute of Electrical and Electronics Engineers Inc.
Additional Information:	The copyright for this article belongs to the Authors.
Keywords:	Audio recordings; Deep neural networks; Gaussian distribution; Speech, Accent identifications; Classification system; Gaussian Mixture Model; I vectors; Network-based modeling; Relative performance; Speech recognition systems; Speech recording, Speech recognition
Department/Centre:	Division of Electrical Sciences > Electrical Engineering
Date Deposited:	01 Sep 2022 04:01
Last Modified:	01 Sep 2022 04:01
URI:	https://eprints.iisc.ac.in/id/eprint/76328

Actions (login required)

View Item


	Powered by EPrints		A service from The J.R.D. Tata Memorial Library Indian Institute of Science, Bengaluru-560012, India