Speaker model selection based on the Bayesian information criterion applied to unsupervised speaker indexing

Nishida, M.; Kawahara, T.

このアイテムのアクセス数: 517

http://hdl.handle.net/2433/128903

このアイテムのファイル:

ファイル	記述	サイズ	フォーマット
TSA.2005.848890.pdf		470.14 kB	Adobe PDF	見る/開く

完全メタデータレコード

DCフィールド	値	言語
dc.contributor.author	Nishida, M.	en
dc.contributor.author	Kawahara, T.	en
dc.contributor.alternative	河原, 達也	ja
dc.date.accessioned	2010-10-21T01:49:58Z	-
dc.date.available	2010-10-21T01:49:58Z	-
dc.date.issued	2005-07	-
dc.identifier.issn	1063-6676	-
dc.identifier.uri	http://hdl.handle.net/2433/128903	-
dc.description.abstract	In conventional speaker recognition tasks, the amount of training data is almost the same for each speaker, and the speaker model structure is uniform and specified manually according to the nature of the task and the available size of the training data. In real-world speech data such as telephone conversations and meetings, however, serious problems arise in applying a uniform model because variations in the utterance durations of speakers are large, with numerous short utterances. We therefore propose a flexible framework in which an optimal speaker model (GMM or VQ) is automatically selected based on the Bayesian Information Criterion (BIC) according to the amount of training data available. The framework makes it possible to use a discrete model when the data is sparse, and to seamlessly switch to a continuous model after a large amount of data is obtained. The proposed framework was implemented in unsupervised speaker indexing of a discussion audio. For a real discussion archive with a total duration of 10 hours, we demonstrate that the proposed method has higher indexing performance than that of conventional methods. The speaker index is also used to adapt a speaker-independent acoustic model to each participant for automatic transcription of the discussion. We demonstrate that speaker indexing with our method is sufficiently accurate for adaptation of the acoustic model.	en
dc.format.mimetype	application/pdf	-
dc.language.iso	eng	-
dc.publisher	IEEE	en
dc.rights	© 2005 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other users, including reprinting/ republishing this material for advertising or promotional purposes, creating new collective works for resale or redistribution to servers or lists, or reuse of any copyrighted components of this work in other works.	en
dc.title	Speaker model selection based on the Bayesian information criterion applied to unsupervised speaker indexing	en
dc.type	journal article	-
dc.type.niitype	Journal Article	-
dc.identifier.ncid	AA10888994	-
dc.identifier.jtitle	IEEE Transactions on Speech and Audio Processing	en
dc.identifier.volume	13	-
dc.identifier.issue	4	-
dc.identifier.spage	583	-
dc.identifier.epage	592	-
dc.relation.doi	10.1109/TSA.2005.848890	-
dc.textversion	publisher	-
dcterms.accessRights	open access	-
出現コレクション:	学術雑誌掲載論文等

アイテムの簡略レコードを表示する

Export to RefWorks

このリポジトリに保管されているアイテムはすべて著作権により保護されています。