• Publications

  • All
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • All
  • Speech Synthesis
  • Speech Recognition
  • Speaker Recognition
  • Speech Signal Processing
  • Affective Computing
  • Multimodal Speech and Language Processing
  • All
  • Journal
  • Conference
Xixin WU, Yuewen CAO, Hui LU, Songxiang LIU, Disong WANG, Zhiyong WU, Xunying LIU, Helen MENG, "Speech Emotion Recognition Using Sequential Capsule Networks," IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), vol. 29, pp. 3280-3291. IEEE, October, 2021. (SCI: WOS:000714713700004, EI: 20214311082562, CCF-A)
Xixin WU, Yuewen CAO, Hui LU, Songxiang LIU, Shiyin KANG, Zhiyong WU, Xunying LIU, Helen MENG, "Exemplar-Based Emotive Speech Synthesis," IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP), vol. 29, pp. 874-886. IEEE, January, 2021. (SCI: WOS:000619310400001, EI: 20210409830187, CCF-A)
Yingmei GUO, Linjun SHOU, Jian PEI, Ming GONG, Mingxing XU, Zhiyong WU and Daxin JIANG, "Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding," [in] Proc. 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1-12. Punta Cana, Dominican Republic, November 7-11, 2021. (EI: 20221411909706, THU-A)
Yaohua BU, Tianyi MA, Weijun LI, Hang ZHOU, Jia JIA, Shengqi CHEN, Kaiyuan XU, Dachuan SHI, Haozhe WU, Zhihan YANG, Kun LI, Zhiyong WU, "PTeacher: A Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback," [in] Proc. 2021 CHI Conference on Human Factors in Computing Systems (CHI), pp. 1-14. Yokohama, Japan, May 8-13, 2021. (EI: 20212210439123, CCF-A)
Suping ZHOU, Jia JIA, Zhiyong WU, Zhihan YANG, Yanfeng WANG, Wei CHEN, Fanbo MENG, Shuo HUANG, Jialie SHEN, Xiaochuan WANG, "Inferring Emotion from Large-Scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach," [in] Proc. the 35th AAAI Conference on Artificial Intelligence (AAAI), pp. 6039-6047. Virtual, Online, February 2-9, 2021. (EI: 20222012114882, CCF-A)
Runnan LI, Zhiyong WU, Jia JIA, Yaohua BU, Sheng ZHAO, Helen MENG, "Towards Discriminative Representation Learning for Speech Emotion Recognition," [in] Proc. International Joint Conference on Artificial Intelligence (IJCAI), pp. 5060-5066. Macao, China, August 10-16, 2019. (EI: 20194607696464, CCF-A)
Yishuang NING, Sheng HE, Zhiyong WU, Chunxiao XING, Liangjie ZHANG, "A Review of Deep Learning Based Speech Synthesis," Applied Sciences-Basel, vol. 9, no. 19, pp. 4050. MDPI, September, 2019. (SCI: WOS:000496258100108)

Runnan LI, Zhiyong WU, Jia JIA, Jingbei LI, Wei CHEN, Helen MENG, "Inferring User Emotive State Changes in Realistic Human-Computer Conversational Dialogs," [in] Proc. ACM Multimedia Conference (ACM MM), pp. 136-144. Seoul, Korea, October 22-26, 2018. (EI: 20185006246269, CCF-A)
Kun LI, Shaoguang MAO, Xu LI, Zhiyong WU, Helen MENG, "Automatic Lexical Stress and Pitch Accent Detection for L2 English Speech using Multi-Distribution Deep Neural Networks," Speech Communication (Speech Com), vol. 96, pp. 28-36. Elsevier, February, 2018. (SCI: WOS:000424723700003, EI: 20174704448303, CCF-B)
Yishuang NING, Jia JIA, Zhiyong WU, Runnan LI, Yongsheng AN, Yanfeng WANG, Helen MENG, "Multi-task Deep Learning for User Intention Understanding in Speech Interaction Systems," [in] Proc. the 31th AAAI Conference on Artificial Intelligence (AAAI), pp. 161-167. San Francisco, USA, February 4-9, 2017. (EI: 20174104242835, CCF-A)
Zhiyong WU, Yishuang NING, Xiao ZANG, Jia JIA, Fanbo MENG, Helen MENG, Lianhong CAI, "Generating Emphatic Speech with Hidden Markov Model for Expressive Speech Synthesis," Multimedia Tools and Applications (MTA), vol. 74, no. 22, pp. 9909-9925. Springer, July, 2015. (SCI: WOS:000364019400005, EI: 20143600027913, CCF-C)
Zhiyong WU, Kai ZHAO, Xixin WU, Xinyu LAN, Helen MENG, "Acoustic to Articulatory Mapping with Deep Neural Network," Multimedia Tools and Applications (MTA), vol. 74, no. 22, pp. 9889-9907. Springer, August, 2015. (SCI: WOS:000364019400004, EI: 20143600014973, CCF-C)

Qi LYU, Zhiyong WU, Jun ZHU, "Polyphonic Music Modelling with LSTM-RTRBM," [in] Proc. ACM Multimedia Conference (ACM MM), pp. 991-994. Brisbane, Australia, October 26-30, 2015. (EI: 20161602252616, CCF-A)

Qi LYU, Zhiyong WU, Jun ZHU, Helen MENG, "Modelling High-dimensional Sequences with LSTM-RTRBM: Application to Polyphonic Music Generation," [in] Proc. International Joint Conference on Artificial Intelligence (IJCAI), pp. 4138-4139. Buenos Aires, Argentina, July 25-31, 2015. (EI: 20155101693661, CCF-A)
Fanbo MENG, Zhiyong WU, Jia JIA, Helen MENG, Lianhong CAI, "Synthesizing English Emphatic Speech for Multimodal Corrective Feedback in Computer-Aided Pronunciation Training," Multimedia Tools and Applications (MTA), vol. 73, no. 1, pp. 463-489. Springer, September, 2014. (SCI: WOS:000342418700022, EI: 20143600046713, CCF-C)
Jia JIA, Zhiyong WU, Shen ZHANG, Helen MENG, Lianhong CAI, "Head and Facial Gestures Synthesis using PAD Model for an Expressive Talking Avatar," Multimedia Tools and Applications (MTA), vol. 73, no. 1, pp. 439-461. Springer, September, 2014. (SCI: WOS:000342418700023, EI: 20143600046670, CCF-C)
Zhiyong WU, Helen M. MENG, Hongwu YANG, Lianhong CAI, "Modeling the Expressivity of Input Text Semantics for Chinese Text-to-Speech Synthesis in a Spoken Dialog System," IEEE Transaction on Audio, Speech and Language Processing (TASLP), vol. 17, no. 8, pp. 1567-1577. IEEE, November, 2009. (SCI: WOS:000268903600010, EI: 20093612281690, CCF-A)