Wav2vec 2.0/HuBERT语音情感识别、演讲人核查和口语理解的精调Wav2vec 2.0/HuBERT基准 (A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding) - 专知论文

会员服务 ·

0

模型评估 · 语音识别 · 可理解性 · Weight · 情景 ·

2022 年 10 月 3 日

A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding

翻译：Wav2vec 2.0/HuBERT语音情感识别、演讲人核查和口语理解的精调Wav2vec 2.0/HuBERT基准

Yingzhi Wang,Abdelmoumene Boumadane,Abdelwahab Heba

from arxiv, 7 pages, 2 figures

Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this work, we explored partial fine-tuning and entire fine-tuning on wav2vec 2.0 and HuBERT pre-trained models for three non-ASR speech tasks: Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding. With simple proposed downstream frameworks, the best scores reached 79.58% weighted accuracy on speaker-dependent setting and 73.01% weighted accuracy on speaker-independent setting for Speech Emotion Recognition on IEMOCAP, 2.36% equal error rate for Speaker Verification on VoxCeleb1, 89.38% accuracy for Intent Classification and 78.92% F1 for Slot Filling on SLURP, showing the strength of fine-tuned wav2vec 2.0 and HuBERT on learning prosodic, voice-print and semantic representations.

翻译：诸如 wav2vec 2. 0 和 HuBERT 等自我监督的演讲模式在自动语音识别方面正在取得革命性的进展。但是,这些模式并没有被完全证明能产生比ASR更好的业绩。在这项工作中,我们探讨了对 wav2vec 2.0 和 HuBERT 三个非ASR 演讲任务进行部分微调和整个微调:语音识别、发言人核查和口语理解等经过培训的模式。在简单提议的下游框架下游框架下,最优得分达到依赖发言设置的加权精度79.58%,对IDEMOCAP 的言语识别设置的加权精度为73.01%。 VoxCeleb1、Inttelecation 准确度为2.36%;SLURP 填充缩放为78.92% F1,显示微调的 wav2vec 2.0 和HuBERT在学习Prosodic、语音和语义表达方面的力量。

0

相关内容

模型评估

机器学习系统设计系统评估标准

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

专知会员服务

68+阅读 · 2022年3月20日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

139+阅读 · 2020年7月10日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

ExBert — 可视化分析Transformer学到的表示

ExBert — 可视化分析Transformer学到的表示

专知会员服务

32+阅读 · 2019年10月16日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

亲水性氨基酸离子液体吸收CO2的传质-反应机理

国家自然科学基金

0+阅读 · 2014年12月31日

天然来源卤酚类高活性衍生物LM49对LPS诱导的血管内皮炎症MAPK信号通路的调控作用与机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

材料表面化学调控血管内皮细胞粘附和迁移的生物力学机制

国家自然科学基金

0+阅读 · 2013年12月31日

Pictet–Spengler类反应机理的理论研究和新反应设计

国家自然科学基金

0+阅读 · 2013年12月31日

PCAF乙酰化修饰XBP1s蛋白对糖尿病小鼠血糖稳态的调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

高迁移率族蛋白1通过上调肝癌病人kupffer细胞Toll样受体和IL-33表达来促进Th17细胞的功能

国家自然科学基金

0+阅读 · 2012年12月31日

纳米颗粒和纳米柱体的力学行为研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向青藏高原的地表微波辐射建模及多年土壤水分反演

国家自然科学基金

0+阅读 · 2012年12月31日

跨语言信息检索中的机器翻译研究

国家自然科学基金

2+阅读 · 2011年12月31日

富含半胱氨酸的酸性分泌蛋白SPARC在胃癌细胞中的表达和调控

国家自然科学基金

0+阅读 · 2009年12月31日

A Diffeomorphic Flow-based Variational Framework for Multi-speaker Emotion Conversion

A Diffeomorphic Flow-based Variational Framework for Multi-speaker Emotion Conversion

Arxiv

0+阅读 · 2022年11月9日

Zero-Label Prompt Selection

Arxiv

0+阅读 · 2022年11月9日

UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language

Arxiv

0+阅读 · 2022年11月8日

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

Arxiv

0+阅读 · 2022年11月8日

Shapes of Emotions: Multimodal Emotion Recognition in Conversations via Emotion Shifts

Shapes of Emotions: Multimodal Emotion Recognition in Conversations via Emotion Shifts

Arxiv

0+阅读 · 2022年11月7日

SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing

Arxiv

0+阅读 · 2022年11月4日

SPEAKER VGG CCT: Cross-corpus Speech Emotion Recognition with Speaker Embedding and Vision Transformers

Arxiv

0+阅读 · 2022年11月4日

Understanding Diffusion Models: A Unified Perspective

Arxiv

14+阅读 · 2022年8月25日

BERT for Joint Intent Classification and Slot Filling

Arxiv

12+阅读 · 2019年2月28日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

73+阅读 · 2018年12月22日

VIP会员

文章信息

相关主题

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

专知会员服务

68+阅读 · 2022年3月20日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

139+阅读 · 2020年7月10日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

ExBert — 可视化分析Transformer学到的表示

ExBert — 可视化分析Transformer学到的表示

专知会员服务

32+阅读 · 2019年10月16日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

【新书发布】原作者MarcG.Bellemare发布315页分布强化学习书籍(DistributionalRL)

深度强化学习实验室

1+阅读 · 2022年1月11日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

A Diffeomorphic Flow-based Variational Framework for Multi-speaker Emotion Conversion

A Diffeomorphic Flow-based Variational Framework for Multi-speaker Emotion Conversion

Arxiv

0+阅读 · 2022年11月9日

Zero-Label Prompt Selection

Arxiv

0+阅读 · 2022年11月9日

UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language

Arxiv

0+阅读 · 2022年11月8日

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

Arxiv

0+阅读 · 2022年11月8日

Shapes of Emotions: Multimodal Emotion Recognition in Conversations via Emotion Shifts

Shapes of Emotions: Multimodal Emotion Recognition in Conversations via Emotion Shifts

Arxiv

0+阅读 · 2022年11月7日

SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing

Arxiv

0+阅读 · 2022年11月4日

SPEAKER VGG CCT: Cross-corpus Speech Emotion Recognition with Speaker Embedding and Vision Transformers

Arxiv

0+阅读 · 2022年11月4日

Understanding Diffusion Models: A Unified Perspective

Arxiv

14+阅读 · 2022年8月25日

BERT for Joint Intent Classification and Slot Filling

Arxiv

12+阅读 · 2019年2月28日

A Survey on Deep Learning for Named Entity Recognition

A Survey on Deep Learning for Named Entity Recognition

Arxiv

73+阅读 · 2018年12月22日

相关基金

亲水性氨基酸离子液体吸收CO2的传质-反应机理

国家自然科学基金

0+阅读 · 2014年12月31日

天然来源卤酚类高活性衍生物LM49对LPS诱导的血管内皮炎症MAPK信号通路的调控作用与机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

材料表面化学调控血管内皮细胞粘附和迁移的生物力学机制

国家自然科学基金

0+阅读 · 2013年12月31日

Pictet–Spengler类反应机理的理论研究和新反应设计

国家自然科学基金

0+阅读 · 2013年12月31日

PCAF乙酰化修饰XBP1s蛋白对糖尿病小鼠血糖稳态的调控研究

国家自然科学基金

0+阅读 · 2012年12月31日

高迁移率族蛋白1通过上调肝癌病人kupffer细胞Toll样受体和IL-33表达来促进Th17细胞的功能

国家自然科学基金

0+阅读 · 2012年12月31日

纳米颗粒和纳米柱体的力学行为研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向青藏高原的地表微波辐射建模及多年土壤水分反演

国家自然科学基金

0+阅读 · 2012年12月31日

跨语言信息检索中的机器翻译研究

国家自然科学基金

2+阅读 · 2011年12月31日

富含半胱氨酸的酸性分泌蛋白SPARC在胃癌细胞中的表达和调控

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员