实现嘈杂回响环境下的实时单通道语音分离 (Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments) - 专知论文

会员服务 ·

0

语音分离 · 混响 · 单通道 · 通道 · 度量 ·

2023 年 4 月 17 日

Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments

翻译：实现嘈杂回响环境下的实时单通道语音分离

Julian Neri,Sebastian Braun

from arxiv, to appear in ICASSP 2023

Real-time single-channel speech separation aims to unmix an audio stream captured from a single microphone that contains multiple people talking at once, environmental noise, and reverberation into multiple de-reverberated and noise-free speech tracks, each track containing only one talker. While large state-of-the-art DNNs can achieve excellent separation from anechoic mixtures of speech, the main challenge is to create compact and causal models that can separate reverberant mixtures at inference time. In this paper, we explore low-complexity, resource-efficient, causal DNN architectures for real-time separation of two or more simultaneous speakers. A cascade of three neural network modules are trained to sequentially perform noise-suppression, separation, and de-reverberation. For comparison, a larger end-to-end model is trained to output two anechoic speech signals directly from noisy reverberant speech mixtures. We propose an efficient single-decoder architecture with subtractive separation for real-time recursive speech separation for two or more speakers. Evaluation on real monophonic recordings of speech mixtures, according to speech separation measures like SI-SDR, perceptual measures like DNS-MOS, and a novel proposed channel separation metric, show that these compact causal models can separate speech mixtures with low latency, and perform on par with large offline state-of-the-art models like SepFormer.

翻译：实时单通道语音分离的目标是将采集自单一麦克风的音频流中含有多人同时说话、环境噪声和混响的内容分离成多个去混响且不带噪声的语音轨道，每个轨道只包含一个说话者。尽管最先进的大型DNN可以从无混响语音信号中实现出色的分离，但主要挑战在于创建紧凑且因果的模型，在推理时可以分离混响信号。在本文中，我们探索了用于实时分离两个或多个同时说话者的低复杂度、资源高效、因果的DNN体系结构，该体系结构由一系列三个神经网络模块级联组成，分别用于顺序执行噪声抑制、分离和去混响。为进行比较，还通过训练较大的端到端模型，直接从含噪混响的语音混合物中输出两个无混响语音信号。我们提出了一种高效的单解码器体系结构，并采用减法分离实现对两个或多个说话者进行递归语音分离的实时操作。根据语音分离度量（例如SI-SDR）、感知度量（例如DNS-MOS）和一种新颖的提出的通道分离度量，对实际的语音混合单声道录音进行评估，表明这些紧凑因果型模型可以实现低延迟的语音分离，性能与大型线下最先进的模型SepFormer相当。

0

相关内容

语音分离

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

专知会员服务

22+阅读 · 2022年7月19日

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【KDD2020】CAST:一种基于相关关系的多尺度数据自适应光谱聚类算法,CAST: A Correlation-based Adaptive Spectral Clustering Algorithm on Multi-scale Data

【KDD2020】CAST:一种基于相关关系的多尺度数据自适应光谱聚类算法,CAST: A Correlation-based Adaptive Spectral Clustering Algorithm on Multi-scale Data

专知会员服务

20+阅读 · 2020年6月11日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【Google Research】Wavesplit:通过说话者聚类实现端到端的语音分离，Wavesplit: End-to-End Speech Separation by Speaker Clustering

【Google Research】Wavesplit:通过说话者聚类实现端到端的语音分离，Wavesplit: End-to-End Speech Separation by Speaker Clustering

专知会员服务

19+阅读 · 2020年2月26日

GeoffreyHinton-ICML2020投稿论文-偏转对抗攻击 Deflecting Adversarial Attacks

GeoffreyHinton-ICML2020投稿论文-偏转对抗攻击 Deflecting Adversarial Attacks

专知会员服务

24+阅读 · 2020年2月22日

Google AI博客解读论文《Reformer: The Efficient Transformer》，百万量级注意力机制

Google AI博客解读论文《Reformer: The Efficient Transformer》，百万量级注意力机制

专知会员服务

70+阅读 · 2020年1月17日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

ECCV 2022 | 底层视觉新任务：Blind Image Decomposition

ECCV 2022 | 底层视觉新任务：Blind Image Decomposition

PaperWeekly

0+阅读 · 2022年9月8日

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

专知

3+阅读 · 2022年7月19日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

深度自进化聚类：Deep Self-Evolution Clustering

深度自进化聚类：Deep Self-Evolution Clustering

我爱读PAMI

15+阅读 · 2019年4月13日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇视频分类相关论文—层次标签推断、知识图谱、CNNs、DAiSEE、表观和关系网络、转移学习

【论文推荐】最新六篇视频分类相关论文—层次标签推断、知识图谱、CNNs、DAiSEE、表观和关系网络、转移学习

专知

13+阅读 · 2018年2月18日

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

专知

15+阅读 · 2018年2月13日

【论文推荐】最新5篇度量学习（Metric Learning）相关论文—人脸验证、BIER、自适应图卷积、注意力机制、单次学习

【论文推荐】最新5篇度量学习（Metric Learning）相关论文—人脸验证、BIER、自适应图卷积、注意力机制、单次学习

专知

17+阅读 · 2018年2月11日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

19+阅读 · 2017年12月17日

基于人类3D视觉感应的2D到3D视频转换关键技术研究

国家自然科学基金

2+阅读 · 2015年12月31日

城市污水深度净化过程中微量有机污染物的识别及去除机制

国家自然科学基金

0+阅读 · 2014年12月31日

基于压缩感知的矢量地理数据水印模型研究

国家自然科学基金

0+阅读 · 2013年12月31日

酸敏感离子通道调控负性记忆的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

实时系统的非剥夺资源共享和分层调度

国家自然科学基金

0+阅读 · 2012年12月31日

动态不确定对抗环境下DDoS 攻击鲁棒检测方法研究

国家自然科学基金

2+阅读 · 2012年12月31日

人工耳蜗植入者汉语普通话音调识别和音乐感知的试验研究

国家自然科学基金

0+阅读 · 2012年12月31日

行车环境听觉模型及声音处理关键技术

国家自然科学基金

0+阅读 · 2011年12月31日

大型天文望远镜状态监控与故障诊断技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

压缩采样框架下的自适应稀疏信号感知与重建

国家自然科学基金

0+阅读 · 2009年12月31日

Efficient Near Maximum-Likelihood Reliability-Based Decoding for Short LDPC Codes

Arxiv

0+阅读 · 2023年6月5日

A Novel Vision Transformer with Residual in Self-attention for Biomedical Image Classification

Arxiv

0+阅读 · 2023年6月2日

Extending the Metaverse: Hyper-Connected Smart Environments with Mixed Reality and the Internet of Things

Arxiv

0+阅读 · 2023年6月1日

Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition

Arxiv

0+阅读 · 2023年6月1日

End-to-End Document Classification and Key Information Extraction using Assignment Optimization

Arxiv

0+阅读 · 2023年6月1日

Spoken Language Identification System for English-Mandarin Code-Switching Child-Directed Speech

Arxiv

0+阅读 · 2023年6月1日

Efficient Near Maximum-Likelihood Efficient Near Maximum-Likelihood Reliability-Based Decoding for Short LDPC Codes

Arxiv

0+阅读 · 2023年6月1日

Learning Task-preferred Inference Routes for Gradient De-conflict in Multi-output DNNs

Arxiv

0+阅读 · 2023年5月31日

Integrating Intelligent Reflecting Surface into Base Station: Architecture, Channel Model, and Passive Reflection Design

Arxiv

0+阅读 · 2023年5月31日

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Arxiv

12+阅读 · 2021年5月30日

VIP会员

文章信息

相关主题

相关VIP内容

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

专知会员服务

22+阅读 · 2022年7月19日

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【KDD2020】CAST:一种基于相关关系的多尺度数据自适应光谱聚类算法,CAST: A Correlation-based Adaptive Spectral Clustering Algorithm on Multi-scale Data

【KDD2020】CAST:一种基于相关关系的多尺度数据自适应光谱聚类算法,CAST: A Correlation-based Adaptive Spectral Clustering Algorithm on Multi-scale Data

专知会员服务

20+阅读 · 2020年6月11日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【Google Research】Wavesplit:通过说话者聚类实现端到端的语音分离，Wavesplit: End-to-End Speech Separation by Speaker Clustering

【Google Research】Wavesplit:通过说话者聚类实现端到端的语音分离，Wavesplit: End-to-End Speech Separation by Speaker Clustering

专知会员服务

19+阅读 · 2020年2月26日

GeoffreyHinton-ICML2020投稿论文-偏转对抗攻击 Deflecting Adversarial Attacks

GeoffreyHinton-ICML2020投稿论文-偏转对抗攻击 Deflecting Adversarial Attacks

专知会员服务

24+阅读 · 2020年2月22日

Google AI博客解读论文《Reformer: The Efficient Transformer》，百万量级注意力机制

Google AI博客解读论文《Reformer: The Efficient Transformer》，百万量级注意力机制

专知会员服务

70+阅读 · 2020年1月17日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

ECCV 2022 | 底层视觉新任务：Blind Image Decomposition

ECCV 2022 | 底层视觉新任务：Blind Image Decomposition

PaperWeekly

0+阅读 · 2022年9月8日

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

Transformer 落地出现 | Next-ViT实现工业TensorRT实时落地，超越ResNet、CSWin

专知

3+阅读 · 2022年7月19日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

深度自进化聚类：Deep Self-Evolution Clustering

深度自进化聚类：Deep Self-Evolution Clustering

我爱读PAMI

15+阅读 · 2019年4月13日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新六篇视频分类相关论文—层次标签推断、知识图谱、CNNs、DAiSEE、表观和关系网络、转移学习

【论文推荐】最新六篇视频分类相关论文—层次标签推断、知识图谱、CNNs、DAiSEE、表观和关系网络、转移学习

专知

13+阅读 · 2018年2月18日

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

【论文推荐】最新7篇条件随机场（CRF）相关论文—图像标注、对抗学习、端到端、注意力机制、三维人体姿态、图像分割、行为分割和识别

专知

15+阅读 · 2018年2月13日

【论文推荐】最新5篇度量学习（Metric Learning）相关论文—人脸验证、BIER、自适应图卷积、注意力机制、单次学习

【论文推荐】最新5篇度量学习（Metric Learning）相关论文—人脸验证、BIER、自适应图卷积、注意力机制、单次学习

专知

17+阅读 · 2018年2月11日

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

【论文推荐】最新5篇图像描述生成（Image Caption）相关论文—情感、注意力机制、遥感图像、序列到序列、深度神经结构

专知

66+阅读 · 2018年1月31日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

19+阅读 · 2017年12月17日

相关论文

Efficient Near Maximum-Likelihood Reliability-Based Decoding for Short LDPC Codes

Arxiv

0+阅读 · 2023年6月5日

A Novel Vision Transformer with Residual in Self-attention for Biomedical Image Classification

Arxiv

0+阅读 · 2023年6月2日

Extending the Metaverse: Hyper-Connected Smart Environments with Mixed Reality and the Internet of Things

Arxiv

0+阅读 · 2023年6月1日

Adaptive Contextual Biasing for Transducer Based Streaming Speech Recognition

Arxiv

0+阅读 · 2023年6月1日

End-to-End Document Classification and Key Information Extraction using Assignment Optimization

Arxiv

0+阅读 · 2023年6月1日

Spoken Language Identification System for English-Mandarin Code-Switching Child-Directed Speech

Arxiv

0+阅读 · 2023年6月1日

Efficient Near Maximum-Likelihood Efficient Near Maximum-Likelihood Reliability-Based Decoding for Short LDPC Codes

Arxiv

0+阅读 · 2023年6月1日

Learning Task-preferred Inference Routes for Gradient De-conflict in Multi-output DNNs

Arxiv

0+阅读 · 2023年5月31日

Integrating Intelligent Reflecting Surface into Base Station: Architecture, Channel Model, and Passive Reflection Design

Arxiv

0+阅读 · 2023年5月31日

SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition

Arxiv

12+阅读 · 2021年5月30日

相关基金

基于人类3D视觉感应的2D到3D视频转换关键技术研究

国家自然科学基金

2+阅读 · 2015年12月31日

城市污水深度净化过程中微量有机污染物的识别及去除机制

国家自然科学基金

0+阅读 · 2014年12月31日

基于压缩感知的矢量地理数据水印模型研究

国家自然科学基金

0+阅读 · 2013年12月31日

酸敏感离子通道调控负性记忆的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

实时系统的非剥夺资源共享和分层调度

国家自然科学基金

0+阅读 · 2012年12月31日

动态不确定对抗环境下DDoS 攻击鲁棒检测方法研究

国家自然科学基金

2+阅读 · 2012年12月31日

人工耳蜗植入者汉语普通话音调识别和音乐感知的试验研究

国家自然科学基金

0+阅读 · 2012年12月31日

行车环境听觉模型及声音处理关键技术

国家自然科学基金

0+阅读 · 2011年12月31日

大型天文望远镜状态监控与故障诊断技术研究

国家自然科学基金

0+阅读 · 2011年12月31日

压缩采样框架下的自适应稀疏信号感知与重建

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员