The Infona portal uses cookies, i.e. strings of text saved by a browser on the user's device. The portal can access those files and use them to remember the user's data, such as their chosen settings (screen view, interface language, etc.), or their login data. By using the Infona portal the user accepts automatic saving and using this information for portal operation purposes. More information on the subject can be found in the Privacy Policy and Terms of Service. By closing this window the user confirms that they have read the information on cookie usage, and they accept the privacy policy and the way cookies are used by the portal. You can change the cookie settings in your browser.
The audio stream is an important component of a sports video. In this paper, we present a system for audio segmentation and classification, which can segment and classify the sports audio stream into speech, non-speech very well. The novel point in our research is that we apply the segmentation and clustering method which is often used in speaker diarization system for broadcast news to the analysis...
Text normalization is an important component in mandarin Text-to-Speech system. This paper develops a taxonomy of Non-Standard Words (NSW's) based on a Large-scale Chinese corpus and proposes a three-stage text normalization strategy: Finite State Automata (FSA) for initial classification, Maximum Entropy (ME) Classifier & Rules for further classification and General Rules for standard word conversion...
For concatenative speech synthesis based on non-uniform unit selection, the key to improve the synthetic quality is the careful designing of measuring criteria respect to the units adopted. With our previous hierarchical non-uniform unit selection framework (Xu et al., 2007), two measurements for selecting optimal non-uniform units during searching at different layers are proposed in this paper, including...
This paper comparatively evaluated various knowledge sources and smoothing algorithms for pronunciation disambiguation in Mandarin TTS (text-to-speech) systems under maximum entropy (maxent) framework. In particular, five kinds of knowledge sources, namely characters and their pronunciations, words, their pronunciations and part-of-speech. together with two smoothing algorithms, i.e. Gaussian prior...
In this paper, we investigate SVM-based speaker verification by location in the space of reference speakers. Speaker location is represented by a vector of log-likelihoods of utterance data given reference speaker models. Channel or session variability in speaker locations due to microphone, acoustic environments etc. would impair verification performance. To reduce such variability, Within-Class...
Set the date range to filter the displayed results. You can set a starting date, ending date or both. You can enter the dates manually or choose them from the calendar.