ESPnet-ST-v2：多功能口语翻译工具包 (ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit) - 专知论文

会员服务 ·

0

工具 · 多功能 · 元模型 · 语音翻译 · 离散 ·

2023 年 4 月 10 日

ESPnet-ST-v2: Multipurpose Spoken Language Translation Toolkit

翻译：ESPnet-ST-v2：多功能口语翻译工具包

Brian Yan,Jiatong Shi,Yun Tang,Hirofumi Inaguma,Yifan Peng,Siddharth Dalmia,Peter Polák,Patrick Fernandes,Dan Berrebbi,Tomoki Hayashi,Xiaohui Zhang,Zhaoheng Ni,Moto Hira,Soumi Maiti,Juan Pino,Shinji Watanabe

ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community. ESPnet-ST-v2 supports 1) offline speech-to-text translation (ST), 2) simultaneous speech-to-text translation (SST), and 3) offline speech-to-speech translation (S2ST) -- each task is supported with a wide variety of approaches, differentiating ESPnet-ST-v2 from other open source spoken language translation toolkits. This toolkit offers state-of-the-art architectures such as transducers, hybrid CTC/attention, multi-decoders with searchable intermediates, time-synchronous blockwise CTC/attention, Translatotron models, and direct discrete unit models. In this paper, we describe the overall design, example models for each task, and performance benchmarking behind ESPnet-ST-v2, which is publicly available at https://github.com/espnet/espnet.

翻译：ESPnet-ST-v2是一款开源工具包，是ESPnet-ST工具包的一次改版，满足了口语翻译社区扩大的需求。ESPnet-ST-v2支持1）离线语音到文本翻译（ST），2）同时语音到文本翻译（SST），以及3）离线语音到语音翻译（S2ST）--每个任务都支持多种方法，这使得ESPnet-ST-v2与其他开源口语翻译工具包有所区别。该工具包提供了最先进的架构，如传输器、混合CTC/注意力、具有可搜索中间结果的多解码器、同步块CTC/注意力、Translatotron模型和直接离散单元模型。在本文中，我们介绍了ESPnet-ST-v2的整体设计、每个任务的示例模型和性能基准测试，该工具包可在https://github.com/espnet/espnet公开获取。

0

相关内容

【李老师400+页的ChatGPT全面介绍PPT】《ChatGPT的前世今生》

【李老师400+页的ChatGPT全面介绍PPT】《ChatGPT的前世今生》

专知会员服务

173+阅读 · 2023年4月13日

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

专知会员服务

44+阅读 · 2020年11月2日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【论文翻译】2020最新预训练语言模型综述：Pre-trained Models for Natural Language Processing: A Survey

【论文翻译】2020最新预训练语言模型综述：Pre-trained Models for Natural Language Processing: A Survey

专知会员服务

94+阅读 · 2020年4月13日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

PyTorch语义分割开源库semseg

PyTorch语义分割开源库semseg

极市平台

25+阅读 · 2019年6月6日

Python中文分词工具大合集：安装、使用和测试

Python中文分词工具大合集：安装、使用和测试

AINLP

11+阅读 · 2019年5月13日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

自然语言处理 | 使用Spacy 进行自然语言处理

自然语言处理 | 使用Spacy 进行自然语言处理

机器学习和数学

19+阅读 · 2018年8月22日

【论文推荐】最新八篇情感分析相关论文—注意力网络、多模态情感分析、情感分析局限性、跨语言情感分类、多语言情感分析

【论文推荐】最新八篇情感分析相关论文—注意力网络、多模态情感分析、情感分析局限性、跨语言情感分类、多语言情感分析

专知

52+阅读 · 2018年6月28日

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

专知

10+阅读 · 2018年6月8日

【论文推荐】最新5篇聊天机器人（Chatbot）相关论文—深度强化学习、社交聊天机器人小冰、对话聊天助手、序列-序列、动态词汇

【论文推荐】最新5篇聊天机器人（Chatbot）相关论文—深度强化学习、社交聊天机器人小冰、对话聊天助手、序列-序列、动态词汇

专知

23+阅读 · 2018年1月30日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

【数据集】新的YELP数据集官方下载

【数据集】新的YELP数据集官方下载

机器学习研究会

16+阅读 · 2017年8月31日

基于MB-OFDM-UWB的煤矿井下无线多媒体传感器网络若干关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

幼儿汉语口语感知特点及神经机制

国家自然科学基金

0+阅读 · 2014年12月31日

无线传感器网络分布式安全时钟同步算法研究

国家自然科学基金

0+阅读 · 2014年12月31日

维、哈、柯跨语言内容过滤关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于汉语话题的句际关系自动分析研究

国家自然科学基金

0+阅读 · 2012年12月31日

100Gbps吞吐率高速高性能LDPC码解码器设计研究

国家自然科学基金

0+阅读 · 2012年12月31日

计算力学基本计算及可视化工具程序包的开发与集成

国家自然科学基金

2+阅读 · 2012年12月31日

跨语言信息检索中的机器翻译研究

国家自然科学基金

2+阅读 · 2011年12月31日

语义计算与理解的资源共享与测评方法

国家自然科学基金

0+阅读 · 2009年12月31日

异步低功耗LDPC解码器设计

国家自然科学基金

0+阅读 · 2009年12月31日

TranSFormer: Slow-Fast Transformer for Machine Translation

Arxiv

0+阅读 · 2023年5月26日

A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target Training

Arxiv

0+阅读 · 2023年5月26日

AMPERE: AMR-Aware Prefix for Generation-Based Event Argument Extraction Model

Arxiv

0+阅读 · 2023年5月26日

S4M: Generating Radiology Reports by A Single Model for Multiple Body Parts

Arxiv

0+阅读 · 2023年5月26日

DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions

Arxiv

0+阅读 · 2023年5月26日

Grounding Language Models to Images for Multimodal Inputs and Outputs

Arxiv

0+阅读 · 2023年5月26日

CARAMEL: A Succinct Read-Only Lookup Table via Compressed Static Functions

Arxiv

0+阅读 · 2023年5月26日

Scaling Data-Constrained Language Models

Arxiv

0+阅读 · 2023年5月25日

Decomposing Complex Queries for Tip-of-the-tongue Retrieval

Arxiv

0+阅读 · 2023年5月24日

DeepSeek: Content Based Image Search & Retrieval

Arxiv

13+阅读 · 2018年1月11日

VIP会员

文章信息

相关主题

相关VIP内容

【李老师400+页的ChatGPT全面介绍PPT】《ChatGPT的前世今生》

【李老师400+页的ChatGPT全面介绍PPT】《ChatGPT的前世今生》

专知会员服务

173+阅读 · 2023年4月13日

自然语言处理顶会NAACL2022最佳论文出炉！

自然语言处理顶会NAACL2022最佳论文出炉！

专知会员服务

43+阅读 · 2022年6月30日

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

纽约大学最新《语音识别Speech Recognition》2020课程，不可错过！

专知会员服务

44+阅读 · 2020年11月2日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【论文翻译】2020最新预训练语言模型综述：Pre-trained Models for Natural Language Processing: A Survey

【论文翻译】2020最新预训练语言模型综述：Pre-trained Models for Natural Language Processing: A Survey

专知会员服务

94+阅读 · 2020年4月13日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津大学博士论文】将序列结构与几何结构融入深度神经网络

工程视角：影响战争进程的小型无人机

企业级AI应用开发：从技术选型到生产落地

AI生成代码缺陷综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

PyTorch语义分割开源库semseg

PyTorch语义分割开源库semseg

极市平台

25+阅读 · 2019年6月6日

Python中文分词工具大合集：安装、使用和测试

Python中文分词工具大合集：安装、使用和测试

AINLP

11+阅读 · 2019年5月13日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

自然语言处理 | 使用Spacy 进行自然语言处理

自然语言处理 | 使用Spacy 进行自然语言处理

机器学习和数学

19+阅读 · 2018年8月22日

【论文推荐】最新八篇情感分析相关论文—注意力网络、多模态情感分析、情感分析局限性、跨语言情感分类、多语言情感分析

【论文推荐】最新八篇情感分析相关论文—注意力网络、多模态情感分析、情感分析局限性、跨语言情感分类、多语言情感分析

专知

52+阅读 · 2018年6月28日

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

专知

10+阅读 · 2018年6月8日

【论文推荐】最新5篇聊天机器人（Chatbot）相关论文—深度强化学习、社交聊天机器人小冰、对话聊天助手、序列-序列、动态词汇

【论文推荐】最新5篇聊天机器人（Chatbot）相关论文—深度强化学习、社交聊天机器人小冰、对话聊天助手、序列-序列、动态词汇

专知

23+阅读 · 2018年1月30日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

【数据集】新的YELP数据集官方下载

【数据集】新的YELP数据集官方下载

机器学习研究会

16+阅读 · 2017年8月31日

相关论文

TranSFormer: Slow-Fast Transformer for Machine Translation

Arxiv

0+阅读 · 2023年5月26日

A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target Training

Arxiv

0+阅读 · 2023年5月26日

AMPERE: AMR-Aware Prefix for Generation-Based Event Argument Extraction Model

Arxiv

0+阅读 · 2023年5月26日

S4M: Generating Radiology Reports by A Single Model for Multiple Body Parts

Arxiv

0+阅读 · 2023年5月26日

DataFinder: Scientific Dataset Recommendation from Natural Language Descriptions

Arxiv

0+阅读 · 2023年5月26日

Grounding Language Models to Images for Multimodal Inputs and Outputs

Arxiv

0+阅读 · 2023年5月26日

CARAMEL: A Succinct Read-Only Lookup Table via Compressed Static Functions

Arxiv

0+阅读 · 2023年5月26日

Scaling Data-Constrained Language Models

Arxiv

0+阅读 · 2023年5月25日

Decomposing Complex Queries for Tip-of-the-tongue Retrieval

Arxiv

0+阅读 · 2023年5月24日

DeepSeek: Content Based Image Search & Retrieval

Arxiv

13+阅读 · 2018年1月11日

相关基金

基于MB-OFDM-UWB的煤矿井下无线多媒体传感器网络若干关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

幼儿汉语口语感知特点及神经机制

国家自然科学基金

0+阅读 · 2014年12月31日

无线传感器网络分布式安全时钟同步算法研究

国家自然科学基金

0+阅读 · 2014年12月31日

维、哈、柯跨语言内容过滤关键技术研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于汉语话题的句际关系自动分析研究

国家自然科学基金

0+阅读 · 2012年12月31日

100Gbps吞吐率高速高性能LDPC码解码器设计研究

国家自然科学基金

0+阅读 · 2012年12月31日

计算力学基本计算及可视化工具程序包的开发与集成

国家自然科学基金

2+阅读 · 2012年12月31日

跨语言信息检索中的机器翻译研究

国家自然科学基金

2+阅读 · 2011年12月31日

语义计算与理解的资源共享与测评方法

国家自然科学基金

0+阅读 · 2009年12月31日

异步低功耗LDPC解码器设计

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员