Abstract for blackburn_icslp96

Proc. ICSLP 96, Philadelphia, October 1996

PSEUDO-ARTICULATORY SPEECH SYNTHESIS FOR RECOGNITION USING AUTOMATIC FEATURE EXTRACTION FROM X-RAY DATA

C.S. Blackburn and S.J. Young

October 1996

We describe a self-organising pseudo-articulatory speech production model (SPM) trained on an X-ray microbeam database, and present results when using the SPM within a speech recognition framework. Given a time-aligned phonemic string, the system uses an explicit statistical model of co-articulation to generate pseudo-articulator trajectories. From these, parametrised speech vectors are synthesised using a set of artificial neural networks (ANNs). We present an analysis of the articulatory information in the database used, and demonstrate the improvements in articulatory modelling accuracy obtained using our co-articulation system. Finally, we give results when using the SPM to re-score N-best utterance transcription lists as produced by the CUED HTK Hidden Markov Model (HMM) speech recognition system. Relative reductions of 18\% in the phoneme error rate and 15\% in the word error rate are achieved.

(ftp:) blackburn_icslp96.ps.gz (http:) blackburn_icslp96.ps.gz
PDF (automatically generated from original PostScript document - may be badly aliased on screen):
(ftp:) blackburn_icslp96.pdf | (http:) blackburn_icslp96.pdf

If you have difficulty viewing files that end '.gz', which are gzip compressed, then you may be able to find tools to uncompress them at the gzip web site.

If you have difficulty viewing files that are in PostScript, (ending '.ps' or '.ps.gz'), then you may be able to find tools to view them at the gsview web site.

We have attempted to provide automatically generated PDF copies of documents for which only PostScript versions have previously been available. These are clearly marked in the database - due to the nature of the automatic conversion process, they are likely to be badly aliased when viewed at default resolution on screen by acroread.