Abstract for tranter_odyssey04

Proc. Odyssey 2004 Speaker and Language Recognition Workshop, June 2004 (Toledo, Spain)


S. E. Tranter and D. A. Reynolds

June 2004

It is often important to be able to automatically label `who spoke when' during some audio data. This paper describes two systems for audio segmentation developed at CUED and MIT-LL and evaluates their performance using the speaker diarisation score defined in the 2003 Rich Transcription Evaluation. A new clustering procedure and BIC-based stopping criterion for the CUED system is introduced which improves both performance and robustness to changes in segmentation. Finally a hybrid `Plug and Play' system is built which combines different parts of the CUED and MIT-LL systems to produce a single system which outperforms both the individual systems.

| (ftp:) tranter_odyssey04.pdf | (http:) tranter_odyssey04.pdf | (ftp:) tranter_odyssey04.ps.gz | (http:) tranter_odyssey04.ps.gz | (http:) tranter_odyssey04.html/ |

If you have difficulty viewing files that end '.gz', which are gzip compressed, then you may be able to find tools to uncompress them at the gzip web site.

If you have difficulty viewing files that are in PostScript, (ending '.ps' or '.ps.gz'), then you may be able to find tools to view them at the gsview web site.

We have attempted to provide automatically generated PDF copies of documents for which only PostScript versions have previously been available. These are clearly marked in the database - due to the nature of the automatic conversion process, they are likely to be badly aliased when viewed at default resolution on screen by acroread.