Yizhong Geng
B.S./M.S. Student in Artificial Intelligence
Beijing University of Posts and Telecommunications
Research Interests
Footprints
Places I have visited.
About
I am currently a B.S./M.S. student in Artificial Intelligence at Beijing University of Posts and Telecommunications, advised by Prof. Ya Li.
My research focuses on text-to-speech, singing voice synthesis, and audio language models, with a particular interest in low-resource and minority-language speech synthesis. I care about the full path from algorithms to products: how to build data, train models, evaluate quality, and polish research prototypes into practical speech systems. My first-author papers have appeared in or been accepted to ICML, ACL, ICME, and INTERSPEECH.
Outside academic research, I also work closely with industry. I was a speech algorithm intern in the Foundation Model team at Li Auto, where I worked on Instruct TTS, continuous-feature-based speech synthesis, and melody-controllable singing voice synthesis for controllable, high-quality in-car voice generation. I am also a co-founder of Beijing Deep Logic Intelligence Technology, where I have helped design and build video direct translation and comic-drama dubbing engines from zero to one for ToB/C product scenarios.
I welcome academic exchange and industry collaboration. Please feel free to contact me.
Selected Publications
View All →Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models
Yizhong Geng, Yanliang Li, Jinghan Yang, Tianhan Jiang, Boxun An, Ya Li, Xiaoyu Shen†
International Conference on Machine Learning (ICML) · 2026
Studies synthetic data scaling for low-resource spoken language models and proposes self-alignment methods that improve stability and expressivity for Thai and Lao speech generation.
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai
Yizhong Geng, Jizhuo Xu, Zeyu Liang, Jinghan Yang, Xiaoyi Shi, Xiaoyu Shen†
ACL · 2025
Builds a scalable industrial Thai TTS framework with data construction, phoneme-tone modeling, and acoustic modeling for low-resource speech synthesis.
Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering
Yizhong Geng, Kecan Mao, Qifei Li, Cong Wang, Yingming Gao, Ruimin Wang, Chunfeng Wang, Hao Li, Ya Li†
Proceedings of INTERSPEECH 2026 · 2026
Stabilizes Instruct-TTS supervision by combining controllable instruction diversification, semantic-drift filtering, and attribute-aligned acoustic perturbation.
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
Yizhong Geng, Wenxin Fu, Kecan Mao, Qifei Li, Yingming Gao, Ruimin Wang, Chunfeng Wang, Hao Li, Ya Li†, Wei Chen†
IEEE International Conference on Multimedia and Expo (ICME) · 2026
Introduces a neural audio codec that embeds melodic structure into discrete representations for high-fidelity and controllable singing voice generation.
News
Instruct-TTS Stabilizer accepted to INTERSPEECH 2026.
Paper on low-resource spoken language models accepted to ICML 2026.
MeloCodec accepted to ICME 2026.
HQ-SVC published at AAAI 2026.
Mel-Refine accepted to ASRU 2025.
EMO-Avatar published at ACM Multimedia 2025.
EEG-based voice conversion work published at INTERSPEECH 2025.
Thai low-resource TTS work published at ACL 2025 Industry Track.