DappetTTS: 一个小型脚印、快速、稳定网络,用于台式机上文字语音语音 (DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech)

With the number of smart devices increasing, the demand for on-device text-to-speech (TTS) increases rapidly. In recent years, many prominent End-to-End TTS methods have been proposed, and have greatly improved the quality of synthesized speech. However, to ensure the qualified speech, most TTS systems depend on large and complex neural network models, and it's hard to deploy these TTS systems on-device. In this paper, a small-footprint, fast, stable network for on-device TTS is proposed, named as DeviceTTS. DeviceTTS makes use of a duration predictor as a bridge between encoder and decoder so as to avoid the problem of words skipping and repeating in Tacotron. As we all know, model size is a key factor for on-device TTS. For DeviceTTS, Deep Feedforward Sequential Memory Network (DFSMN) is used as the basic component. Moreover, to speed up inference, mix-resolution decoder is proposed for balance the inference speed and speech quality. Experiences are done with WORLD and LPCNet vocoder. Finally, with only 1.4 million model parameters and 0.099 GFLOPS, DeviceTTS achieves comparable performance with Tacotron and FastSpeech. As far as we know, the DeviceTTS can meet the needs of most of the devices in practical application.

翻译：智能设备数量不断增加,对智能设备的需求迅速增加。近年来,提出了许多著名的端到端语音技术方法,并大大提高了合成语音的质量。然而,为了确保语言合格,大多数TTS系统依赖于大型和复杂的神经网络模型,很难将这些TTS系统安装在设计上。本文提出了一个小脚印、快速、稳定的TTTS网络,称为“设备TTS”。设备TTS使用一个期限预测器,作为编码器和解码器之间的桥梁,以避免在塔科特罗出现跳过和重复的词的问题。我们都知道,模型大小是TTTS系统使用大型和复杂的神经网络模型的一个关键因素。对于设备TTTS系统,将深喂向前序列记忆网络(DFSMN)用作基本组成部分。此外,为了加快判断,建议混合解析器能够平衡语音模型和语音质量之间的平衡。AsmalS和LPCSFATS最接近性能标准,AsldS 与GPTTS 和SFTFCS 系统最相近40的运行标准。

相关内容

语音合成

关注 491

语音合成（Speech Synthesis），也称为文语转换（Text-to-Speech, TTS,它是将任意的输入文本转换成自然流畅的语音输出。语音合成涉及到人工智能、心理学、声学、语言学、数字信号处理、计算机科学等多个学科技术，是信息处理领域中的一项前沿技术。随着计算机技术的不断提高，语音合成技术从早期的共振峰合成,逐步发展为波形拼接合成和统计参数语音合成，再发展到混合语音合成；合成语音的质量、自然度已经得到明显提高，基本能满足一些特定场合的应用需求。目前，语音合成技术在银行、医院等的信息播报系统、汽车导航系统、自动应答呼叫中心等都有广泛应用，取得了巨大的经济效益。另外，随着智能手机、MP3、PDA 等与我们生活密切相关的媒介的大量涌现，语音合成的应用也在逐渐向娱乐、语音教学、康复治疗等领域深入。可以说语音合成正在影响着人们生活的方方面面。

深度学习搜索，Exploring Deep Learning for Search

专知会员服务

61+阅读 · 2020年5月9日

【论文翻译】NLP注意力机制综述论文翻译，Attention, please! A Critical Review of Neural Attention Models in Natural Language Processing

专知会员服务

96+阅读 · 2020年4月18日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

【ICLR-2020】网络反卷积，NETWORK DECONVOLUTION

专知会员服务

39+阅读 · 2020年2月21日