TTS - VLSP 2021: Development of Smartcall Vietnamese Text-to-Speech

Nguyen Quoc Bao; Le Ba Hoai; Nguyen Van Hoc; Dam Ba Quyen; Nguyen Thu Phuong

doi:10.25073/2588-1086/vnucsce.348

Nguyen Quoc Bao, Le Ba Hoai, Nguyen Van Hoc, Dam Ba Quyen, Nguyen Thu Phuong

PDF

Published Jun 30, 2022

DOI: https://doi.org/10.25073/2588-1086/vnucsce.348

How to Cite

QUOC BAO, Nguyen et al. TTS - VLSP 2021: Development of Smartcall Vietnamese Text-to-Speech. VNU Journal of Science: Computer Science and Communication Engineering, [S.l.], v. 38, n. 1, june 2022. ISSN 2588-1086. Available at: <//jcsce.vnu.edu.vn/index.php/jcsce/article/view/348>. Date accessed: 24 july 2026. doi: https://doi.org/10.25073/2588-1086/vnucsce.348.

ABNT APA BibTeX CBE EndNote - EndNote format (Macintosh & Windows) MLA ProCite - RIS format (Macintosh & Windows) RefWorks Reference Manager - RIS format (Windows only) Turabian

Issue

Vol 38 No 1: Special Issue: The 8th International Workshop on Vietnamese Language and Speech Processing (VLSP 2021)

Section

Special Issue on Vietnamese Language and Speech Processing (VLSP2021)

Abstract

Recent advances in deep learning facilitate the development of end-to-end Vietnamese text-to-speech (TTS) systems with high intelligibility and naturalness in the presence of a clean training corpus. Given a rich source of audio recording data on the Internet, TTS has excellent potential for growth if it can take advantage of this data source. However, the quality of these data is often not sufficient for training TTS systems, e.g., noisy audio. In this paper, we propose an approach that preprocesses noisy found data on the Internet and trains a high-quality TTS model on the processed data. The VLSP-provided training data was thoroughly preprocessed using 1) voice activity detection, 2) automatic speech recognition-based prosodic punctuation insertion, and 3) Spleeter, source separation tool, for separating voice from background music. Moreover, we utilize a state-of-the-art TTS system that takes advantage of the Conditional Variational Autoencoder with the Adversarial Learning model. Our experiment showed that the proposed TTS system trained on the preprocessed data achieved a good result on the provided noisy dataset.

Article Sidebar

Article Details

Main Article Content

Abstract