ASR2K: 约2000年无音频语言的语音识别 (ASR2K: Speech Recognition for Around 2000 Languages without Audio) - 专知论文

会员服务 ·

0

N元 · 统计量 · 语音识别 · MoDELS · 数据集 ·

2022 年 9 月 6 日

ASR2K: Speech Recognition for Around 2000 Languages without Audio

翻译：ASR2K: 约2000年无音频语言的语音识别

Xinjian Li,Florian Metze,David R Mortensen,Alan W Black,Shinji Watanabe

from arxiv, INTERSPEECH 2022

Most recent speech recognition models rely on large supervised datasets, which are unavailable for many low-resource languages. In this work, we present a speech recognition pipeline that does not require any audio for the target language. The only assumption is that we have access to raw text datasets or a set of n-gram statistics. Our speech pipeline consists of three components: acoustic, pronunciation, and language models. Unlike the standard pipeline, our acoustic and pronunciation models use multilingual models without any supervision. The language model is built using n-gram statistics or the raw text dataset. We build speech recognition for 1909 languages by combining it with Crubadan: a large endangered languages n-gram database. Furthermore, we test our approach on 129 languages across two datasets: Common Voice and CMU Wilderness dataset. We achieve 50% CER and 74% WER on the Wilderness dataset with Crubadan statistics only and improve them to 45% CER and 69% WER when using 10000 raw text utterances.

翻译：最新的语音识别模型依靠大型监管数据集, 许多低资源语言都无法获得这些数据。在这项工作中, 我们提出了一个语音识别管道, 不需要目标语言的任何音频。唯一的假设是, 我们能够获得原始文本数据集或一组 n- gram 统计数据。我们的语音验证包括三个组成部分: 音频、发音和语言模型。与标准管道不同, 我们的声频和发音模型使用多种语言模型, 没有任何监督。语言模型是使用 n- gram 统计数据或原始文本数据集构建的。我们通过将它与 Crubadan 合并, 建立1909 语言的语音识别管道: 一个大型濒危语言 n- gram 数据库。此外, 我们测试了我们跨越两个数据集的129种语言: 通用语音和 CMU Warderness 数据集。我们用Crubadan 统计数据在Wilderness 数据集上实现了50% CER 和 74% WER 。我们只用 Crubadan 统计数据将其提高到45% CER 和 69% WER 。

0

相关内容

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

专知会员服务

44+阅读 · 2020年11月2日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

补肾方调节MAVS介导的信号通路发挥抗炎、抗病毒作用机制及其主体疗效中药组

国家自然科学基金

0+阅读 · 2014年12月31日

基于稀疏贝叶斯学习的稳健空时自适应处理研究

国家自然科学基金

2+阅读 · 2013年12月31日

应力对FeRh薄膜磁卡效应的调控研究

国家自然科学基金

0+阅读 · 2013年12月31日

Pictet–Spengler类反应机理的理论研究和新反应设计

国家自然科学基金

0+阅读 · 2013年12月31日

语音评价中的韵律建模和评价方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

激光脉冲湍流大气瞬态传输机理及对FSO通信的影响研究

国家自然科学基金

0+阅读 · 2012年12月31日

雌激素通过ERα介导lncRNA 1200076调节卵巢ERα（+）细胞生物学行为

国家自然科学基金

0+阅读 · 2012年12月31日

时间分辨光谱研究Nitrenium离子与DNA形成致癌加合物的反应机理

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米尺度下Beta钛合金单晶力学行为及变形机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

改进的Unscented卡尔曼滤波与电池组SOC快速精确估计

国家自然科学基金

0+阅读 · 2008年12月31日

Pre-trained Sentence Embeddings for Implicit Discourse Relation Classification

Arxiv

0+阅读 · 2022年10月20日

Large-scale learning of generalised representations for speaker recognition

Arxiv

0+阅读 · 2022年10月20日

Learning to Discover and Detect Objects

Arxiv

0+阅读 · 2022年10月19日

Deep-based quality assessment of medical images through domain adaptation

Arxiv

0+阅读 · 2022年10月19日

Simple and Effective Unsupervised Speech Translation

Arxiv

0+阅读 · 2022年10月18日

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Arxiv

0+阅读 · 2022年10月18日

Perceptual Grouping in Vision-Language Models

Perceptual Grouping in Vision-Language Models

Arxiv

0+阅读 · 2022年10月18日

Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

Arxiv

0+阅读 · 2022年10月18日

Personalization of CTC Speech Recognition Models

Arxiv

0+阅读 · 2022年10月18日

Zero-Shot Transfer Learning for Event Extraction

Arxiv

10+阅读 · 2017年7月4日

VIP会员

文章信息

相关主题

相关VIP内容

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

专知会员服务

44+阅读 · 2020年11月2日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

Deep Research（深度研究）：系统性综述

《革新战术战场空间能力：反无人机系统》报告

【普林斯顿博士论文】用于语音的生成式通用模型

螺旋式开发作为战略资产：美军启示

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

相关论文

Pre-trained Sentence Embeddings for Implicit Discourse Relation Classification

Arxiv

0+阅读 · 2022年10月20日

Large-scale learning of generalised representations for speaker recognition

Arxiv

0+阅读 · 2022年10月20日

Learning to Discover and Detect Objects

Arxiv

0+阅读 · 2022年10月19日

Deep-based quality assessment of medical images through domain adaptation

Arxiv

0+阅读 · 2022年10月19日

Simple and Effective Unsupervised Speech Translation

Arxiv

0+阅读 · 2022年10月18日

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Arxiv

0+阅读 · 2022年10月18日

Perceptual Grouping in Vision-Language Models

Perceptual Grouping in Vision-Language Models

Arxiv

0+阅读 · 2022年10月18日

Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

Arxiv

0+阅读 · 2022年10月18日

Personalization of CTC Speech Recognition Models

Arxiv

0+阅读 · 2022年10月18日

Zero-Shot Transfer Learning for Event Extraction

Arxiv

10+阅读 · 2017年7月4日

相关基金

补肾方调节MAVS介导的信号通路发挥抗炎、抗病毒作用机制及其主体疗效中药组

国家自然科学基金

0+阅读 · 2014年12月31日

基于稀疏贝叶斯学习的稳健空时自适应处理研究

国家自然科学基金

2+阅读 · 2013年12月31日

应力对FeRh薄膜磁卡效应的调控研究

国家自然科学基金

0+阅读 · 2013年12月31日

Pictet–Spengler类反应机理的理论研究和新反应设计

国家自然科学基金

0+阅读 · 2013年12月31日

语音评价中的韵律建模和评价方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

激光脉冲湍流大气瞬态传输机理及对FSO通信的影响研究

国家自然科学基金

0+阅读 · 2012年12月31日

雌激素通过ERα介导lncRNA 1200076调节卵巢ERα（+）细胞生物学行为

国家自然科学基金

0+阅读 · 2012年12月31日

时间分辨光谱研究Nitrenium离子与DNA形成致癌加合物的反应机理

国家自然科学基金

0+阅读 · 2012年12月31日

亚微米尺度下Beta钛合金单晶力学行为及变形机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

改进的Unscented卡尔曼滤波与电池组SOC快速精确估计

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员