ICASSP 2022多渠道多党会议多党会议分配挑战 (Royalflush Speaker Diarization System for ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge)

This paper describes the Royalflush speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription Challenge(M2MeT). Our system comprises speech enhancement, overlapped speech detection, speaker embedding extraction, speaker clustering, speech separation and system fusion. In this system, we made three contributions. First, we propose an architecture of combining the multi-channel and U-Net-based models, aiming at utilizing the benefits of these two individual architectures, for far-field overlapped speech detection. Second, in order to use overlapped speech detection model to help speaker diarization, a speech separation based overlapped speech handling approach, in which the speaker verification technique is further applied, is proposed. Third, we explore three speaker embedding methods, and obtained the state-of-the-art performance on the CNCeleb-E test set. With these proposals, our best individual system significantly reduces DER from 15.25% to 6.40%, and the fusion of four systems finally achieves a DER of 6.30% on the far-field Alimeeting evaluation set.

翻译：本文介绍了提交多频道多党会议分流挑战(M2MET)的皇家脸红色扬声器分化系统。我们的系统包括语音增强、语音探测重叠、语音嵌入提取、发言者群集、语音分离和系统融合。在这个系统中,我们做出了三项贡献。首先,我们提出了一个将多频道和基于U-Net的模型相结合的结构,目的是利用这两个单个结构的好处,进行远方重叠语音检测。第二,为了利用重叠语音检测模型帮助发言者分化,提出了基于语音分解的重叠语音处理方法,进一步应用了语音校验技术。第三,我们探索了三个发言者嵌入方法,并在CNCeleb-E测试集上获得了最先进的表现。有了这些提议,我们最好的单个系统将DER从15.25%大幅降至6.40%,而四个系统的组合最终在远场Alimeeting评价集中实现了6.30%的DER。

相关内容

ICASSP

关注 4

ICASSP是全球最大，最全面的技术会议，重点是信号处理及其应用。会议主题包括但不限于以下主题：音频和声音信号处理、量子信号处理、生物医学信号与图像处理、遥感与信号处理、压缩感知，采样和字典学习、传感器阵列和多通道信号处理、信号处理的设计与实现、大数据信号处理、财务信号处理。官网地址：http://dblp.uni-trier.de/db/conf/icassp/

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

【北邮-腾讯AI】自监督学习音视觉说话人认证，Self-supervised learning for audio-visual speaker diarization

专知会员服务

26+阅读 · 2020年2月16日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【中科院自动化所】序列到序列语音识别的无监督预训练（Unsupervised pre-training for sequence to sequence speech recognition）

专知会员服务

33+阅读 · 2020年1月5日