This study explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike data-intensive end-to-end models, FUL offers a compact, interpretable feature set grounded in phonological principles, enabling scalable and equitable TTS development for low-resource languages. We provide a mapping from language-specific phones to FUL feature vectors via a SAMPA intermediate and incorporate these features into a modified FastSpeech architecture. Experiments were conducted to evaluate their ability to generate native, non-native, and code-mixed speech in English and Mandarin. We ran an experiment with a small dataset and one with a larger dataset, which showed that TTS with FUL features as input could produce intelligible native speech with as little as 8 hours of training data; with 100 hours of training data, intelligible speech could be generated for a language not present in the training data. The approach further supports code-mixed synthesis while preserving consistent timbre and interpretable phonetic control. These results highlight the potential of theory-driven representations for building efficient, scalable, and linguistically informed TTS systems, demonstrating that phonological features can function as both analytical tools and practical inputs for speech technology.
This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike data-intensive end-to-end models, FUL offers a compact, interpretable feature set grounded in phonological principles, enabling scalable and equitable TTS development for low-resource languages. We provide a mapping from language-specific phones to FUL feature vectors via a SAMPA intermediate and incorporate these features into a modified FastSpeech architecture. Experiments were conducted to evaluate their ability to generate native, non-native, and code-mixed speech in English and Mandarin. We ran an experiment with a small dataset and one with a larger dataset, which showed that TTS with FUL features as input could produce intelligible native speech with as little as 8 hours of training data; with 100 hours of training data, intelligible speech could be generated for a language not present in the training data. The approach further supports code-mixed synthesis while preserving consistent timbre and interpretable phonetic control. These results highlight the potential of theory-driven representations for building efficient, scalable, and linguistically informed TTS systems, demonstrating that phonological features can function as both analytical tools and practical inputs for speech technology.
@article{zhang2026featureTTS,author={Zhang, Cong and Zeng, Huinan and Liu, Huang and Zheng, Jiewen},title={{Integrating Human Linguistic Insights into AI: Theory-Driven Representation for Multilingual Text-to-Speech}},journal={Phonetica},year={2026},}
Related talks:
speech tech
Featurally Underspecified Lexicon (FUL) Model: Evidence from multilingual Text-to-Speech
Cong Zhang, Huinan Zeng, Huang Liu, and Jiewen Zheng
The 18th Conference on Laboratory Phonology Online, 23-25 jun 2022
This study investigates whether the phonological features derived from the Featurally Underspecified Lexicon model can be applied in text-to-speech systems to generate native and non-native speech in English and Mandarin. The results supported that phonological features could be used as a feasible input system for languages in or not in the train data. The results lend support to FUL by presenting successfully synthesised output, and by having the output carrying a source-language accent when synthesising a language not in the training data. The TTS process stimulated second language acquisition process and thus also confirm FUL’s ability to account for acquisition.
@conference{zhang2022featurally-c,author={Zhang, Cong and Zeng, Huinan and Liu, Huang and Zheng, Jiewen},title={Featurally Underspecified Lexicon (FUL) Model: Evidence from multilingual Text-to-Speech},booktitle={The 18th Conference on Laboratory Phonology},year={2022},month={23-25 Jun},location={Online}}
Phone & word alignments for 1300 hours of open-source Mandarin speech datasets. Automatically aligned with our own Charsiu Forced Aligner.
@misc{zhang2021TTS,author={Zhang, Cong and Zeng, Huinan},title={{Phonological feature mapping for FeatureTTS}},year={2021},category={dataset},doi={10.5281/ZENODO.5553685}}