First Broadcast News Transcription System for Khmer Language

Fiche du document

Date

2008

Type de document
Périmètre
Langue
Identifiants
Collection

Archives ouvertes

Licences

http://hal.archives-ouvertes.fr/licences/copyright/ , info:eu-repo/semantics/OpenAccess


Mots-clés En

Speech


Citer ce document

Sopheap Seng et al., « First Broadcast News Transcription System for Khmer Language », HAL-SHS : sciences de l'information, de la communication et des bibliothèques, ID : 10670/1.sh940l


Métriques


Partage / Export

Résumé En

In this paper we present an overview on the development of a large vocabulary continuous speech recognition (LVCSR) system for Khmer, the official language of Cambodia, spoken by more than 15 million people. As an under-resourced language, develop a LVCSR system for Khmer is a challenging task. We describe our methodologies for quick language data collection and processing for language modeling and acoustic modeling. For language modeling, we investigate the use of word and sub-word as basic modeling unit in order to see the potential of sub-word units in the case of unsegmented language like Khmer. Grapheme-based acoustic modeling is used to quickly build our Khmer language acoustic model. Furthermore, the approaches and tools used for the development of our system are documented and made publicly available on the web. We hope this will contribute to accelerate the development of LVCSR system for a new language, especially for under-resource languages of developing countries where resources and expertise are limited.

document thumbnail

Par les mêmes auteurs

Sur les mêmes sujets

Sur les mêmes disciplines

Exporter en