根据会议设想(V1)为多目标发言者同时发言 (Simultaneous Speech Extraction for Multiple Target Speakers under the Meeting Scenarios(V1)) - 专知论文

会员服务 ·

0

分离的 · 估计/估计量 · 基准 · 掩码 · 相互独立的 ·

2022 年 6 月 17 日

Simultaneous Speech Extraction for Multiple Target Speakers under the Meeting Scenarios(V1)

翻译：根据会议设想(V1)为多目标发言者同时发言

Bang Zeng,Weiqing Wang,Yuanyuan Bao,Ming Li

from arxiv, 4pages, 3 figures

Recently, the target speech separation or extraction techniques under the meeting scenario have become a hot research trend. We propose a speaker diarization aware multiple target speech separation system (SD-MTSS) to simultaneously extract the voice of each speaker from the mixed speech, rather than requiring a succession of independent processes as presented in previous solutions. SD-MTSS consists of a speaker diarization (SD) module and a multiple target speech separation (MTSS) module. The former one infers the target speaker voice activity detection (TSVAD) states of the mixture, as well as gets different speakers' single-talker audio segments as the reference speech. The latter one employs both the mixed audio and reference speech as inputs, and then it generates an estimated mask. By exploiting the TSVAD decision and the estimated mask, our SD-MTSS model can extract the speech of each speaker concurrently in a conversion recording without additional enrollment audio in advance.Experimental results show that our MTSS model outperforms our baselines with a large margin, achieving 1.38dB SDR, 1.34dB SI-SNR, and 0.13 PESQ improvements over the state-of-the-art SpEx+ baseline on the WSJ0-2mix-extr dataset, respectively. The SD-MTSS system makes a significant improvement than the baseline on the Alimeeting dataset as well.

翻译：最近,会议设想情景下的目标语音分离或提取技术已成为一个热研究趋势。我们建议让发言者分解了解多目标语音分离系统(SD-MTSS),以同时从混合演讲中提取每个发言者的声音,而不是像先前解决方案中显示的那样需要一系列独立程序。SD-MTSS模式由一个发言者分解模块和一个多目标语音分离模块组成。前一位推断出,目标发言者对混合物的语音活动检测(TSVAD)状态,并获得不同发言者的单讲音部分作为参考演讲。后一位使用混合音频和参考演讲作为投入,然后生成一个估计面罩。通过利用TSVAD决定和估计面罩,我们的SD-MSS模型可以同时在转换记录中提取每个发言者的语音,而无需事先增加录制音频谱。显性结果显示,我们的MSS模型比我们的基线大差,实现了1.38dB特别提款权、1.34dB SI-SNR和0.13 PESQ改进了节音频。通过利用TVA决定和估计面系统,使SD-M-MIS的基线比S-MS-MS-MS-MS-S-S-S-SB改进了州基准系统改进。

0

相关内容

分离的

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

“共核”法制备高Ms锌铁氧体纳米颗粒的形成机理

国家自然科学基金

0+阅读 · 2015年12月31日

水平井随钻阵列感应测井环境动态建模及参数估计

国家自然科学基金

0+阅读 · 2014年12月31日

Nrf2-Keap1-ARE信号通路在脊髓损伤后胶质瘢痕形成中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

地基InSAR高边坡三维变形提取方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于可活化胶原蛋白结合肽的肿瘤MMP-14酶活性核素显像研究

国家自然科学基金

0+阅读 · 2012年12月31日

Skutterudite/AgSbTe2系纳米复合热电材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

智能化高分辨太阳耀斑观测终端的研制

国家自然科学基金

0+阅读 · 2009年12月31日

Unscented卡尔曼滤波算法及其在通信中的应用

国家自然科学基金

0+阅读 · 2008年12月31日

磁性Pickering乳液界面流变学研究

国家自然科学基金

0+阅读 · 2008年12月31日

Using soft maximin for risk averse multi-objective decision-making

Using soft maximin for risk averse multi-objective decision-making

Arxiv

0+阅读 · 2022年8月8日

A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation

Arxiv

0+阅读 · 2022年8月8日

Abstractive Meeting Summarization: A Survey

Arxiv

0+阅读 · 2022年8月8日

An Arbitrary High Order and Positivity Preserving Method for the Shallow Water Equations

Arxiv

0+阅读 · 2022年8月8日

DeepGen: Diverse Search Ad Generation and Real-Time Customization

Arxiv

0+阅读 · 2022年8月6日

Parameter Averaging for Robust Explainability

Parameter Averaging for Robust Explainability

Arxiv

0+阅读 · 2022年8月5日

A Masked Image Reconstruction Network for Document-level Relation Extraction

Arxiv

0+阅读 · 2022年8月4日

Risk-Aware Linear Bandits: Theory and Applications in Smart Order Routing

Arxiv

0+阅读 · 2022年8月4日

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

Arxiv

17+阅读 · 2020年6月2日

Event Extraction with Generative Adversarial Imitation Learning

Arxiv

13+阅读 · 2018年4月21日

VIP会员

文章信息

相关主题

估计/估计量

相互独立的

相关VIP内容

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

大语言模型中的检索与结构化增强生成综述

《实现多层防御多轮交战机制的扩展型随机齐射模型》2025年最新83页

【CMU博士论文】交互驱动的人体动作估计与生成

如何避免生成式人工智能在作战中失控失效

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

相关论文

Using soft maximin for risk averse multi-objective decision-making

Using soft maximin for risk averse multi-objective decision-making

Arxiv

0+阅读 · 2022年8月8日

A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation

Arxiv

0+阅读 · 2022年8月8日

Abstractive Meeting Summarization: A Survey

Arxiv

0+阅读 · 2022年8月8日

An Arbitrary High Order and Positivity Preserving Method for the Shallow Water Equations

Arxiv

0+阅读 · 2022年8月8日

DeepGen: Diverse Search Ad Generation and Real-Time Customization

Arxiv

0+阅读 · 2022年8月6日

Parameter Averaging for Robust Explainability

Parameter Averaging for Robust Explainability

Arxiv

0+阅读 · 2022年8月5日

A Masked Image Reconstruction Network for Document-level Relation Extraction

Arxiv

0+阅读 · 2022年8月4日

Risk-Aware Linear Bandits: Theory and Applications in Smart Order Routing

Arxiv

0+阅读 · 2022年8月4日

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

Arxiv

17+阅读 · 2020年6月2日

Event Extraction with Generative Adversarial Imitation Learning

Arxiv

13+阅读 · 2018年4月21日

相关基金

“共核”法制备高Ms锌铁氧体纳米颗粒的形成机理

国家自然科学基金

0+阅读 · 2015年12月31日

水平井随钻阵列感应测井环境动态建模及参数估计

国家自然科学基金

0+阅读 · 2014年12月31日

Nrf2-Keap1-ARE信号通路在脊髓损伤后胶质瘢痕形成中的作用及机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

地基InSAR高边坡三维变形提取方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于可活化胶原蛋白结合肽的肿瘤MMP-14酶活性核素显像研究

国家自然科学基金

0+阅读 · 2012年12月31日

Skutterudite/AgSbTe2系纳米复合热电材料研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

智能化高分辨太阳耀斑观测终端的研制

国家自然科学基金

0+阅读 · 2009年12月31日

Unscented卡尔曼滤波算法及其在通信中的应用

国家自然科学基金

0+阅读 · 2008年12月31日

磁性Pickering乳液界面流变学研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员