The page
Two columns, a numbered figure, a hanging indent. Nobody set any of it.
Text-to-speech synthesis for low-resource languages
U. Jha, P. Adhikari, P. R. Acharya
Department of Electronics and Computer Engineering
Abstract. This paper presents a synthesis system for a language with little recorded speech available, trained on a corpus assembled from public readings and evaluated against native listeners.
Index Terms. speech synthesis, low-resource languages, evaluation
I. Introduction
Speech interfaces remain out of reach for most of the world's languages, because the recorded corpora they are trained on do not exist. Recent work has narrowed the amount of audio required [1].
A. Corpus
Twelve hours of read speech were segmented and aligned, then held out by speaker so that no voice appeared in both the training and evaluation sets.
Fig. 1. Mean opinion score against hours of training audio
Scores rose steeply to eight hours and then flattened, which is consistent with the behaviour reported for comparable systems [2].
B. Alignment
Forced alignment was run twice, once on the raw readings and once after silence trimming, and the second pass recovered a further four per cent of usable segments.
II. Evaluation
Thirty-two listeners rated a randomised set of utterances on a five-point scale. Each listener heard both synthesised and recorded speech without being told which was which, and no utterance was rated twice by the same person.
Agreement between listeners was moderate, and the ordering of systems was stable across every subset we resampled, so the ranking does not depend on which listeners happened to take part.
References
[1] J. Shen et al., “Natural TTS synthesis,” Proc. ICASSP, 2018, pp. 4779–4783.
1
