建设和整理具有多样性意识的语言科学和技术对话社团 (Building and curating conversational corpora for diversity-aware language science and technology) - 专知论文

会员服务 ·

0

INTERACT · 多样性 · 讲稿 · 迹 · CASE ·

2022 年 5 月 10 日

Building and curating conversational corpora for diversity-aware language science and technology

翻译：建设和整理具有多样性意识的语言科学和技术对话社团

Andreas Liesenfeld,Mark Dingemanse

We present an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. Surveying language documentation corpora and other resources that cover 67 languages and varieties from 28 phyla, we describe the compilation and curation process, specify minimal properties of a unified format for interactional data, and develop methods for quality control that take into account turn-taking and timing. Two case studies show the broad utility of conversational data for (i) charting human interactional infrastructure and (ii) tracing challenges and opportunities for current ASR solutions. Linguistically diverse conversational corpora can provide new insights for the language sciences and stronger empirical foundations for language technology.

翻译：我们提出以多种语言建立和整理日常对话社团的分析编审流程和最佳做法准则,调查语言文献社团和其他资源,涵盖来自28个phyla的67种语言和品种,我们描述汇编和整理过程,具体说明互动数据统一格式的最小特性,并制订考虑到考量和时机的质量控制方法,两个案例研究表明,对话数据在(一) 绘制人类互动基础设施图和(二) 追踪当前ASR解决方案的挑战和机遇方面具有广泛效用。语言多样性谈话社团可以为语言科学提供新的洞察力,并为语言技术提供更强有力的经验基础。

0

相关内容

INTERACT

IFIP TC13 Conference on Human-Computer Interaction是人机交互领域的研究者和实践者展示其工作的重要平台。多年来，这些会议吸引了来自几个国家和文化的研究人员。官网链接：http://interact2019.org/

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

二层平衡问题的适定性与算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

对称性破缺条件下耦合系统chimera态的特性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Corin介导的ANP活化在动脉粥样硬化形成及其炎症反应中的作用与机制

国家自然科学基金

0+阅读 · 2012年12月31日

Z-pinch内爆等离子体二维高温辐射磁流体动力学方程及其动力系统研究

国家自然科学基金

0+阅读 · 2009年12月31日

冲击与火作用下弹塑性梁的动力响应及LTB研究

国家自然科学基金

0+阅读 · 2008年12月31日

A Model-Based Approach for Specifying Changes in Replications of Empirical Studies in Computer Science

Arxiv

0+阅读 · 2022年6月27日

Towards Blockchain-Based Secure Data Management for Remote Patient Monitoring

Arxiv

0+阅读 · 2022年6月26日

Visual Auditor: Interactive Visualization for Detection and Summarization of Model Biases

Arxiv

0+阅读 · 2022年6月25日

Trustworthy AI: From Principles to Practices

Arxiv

46+阅读 · 2021年10月4日

Advances and Challenges in Conversational Recommender Systems: A Survey

Arxiv

14+阅读 · 2021年1月23日

VIP会员

文章信息

相关主题

相关VIP内容

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

【斯坦福博士论文】数据、决策与过度依赖：构建可信人工智能的核心挑战

《多域时代中维持弹性军事训练：挑战与机遇》

【AAAI2026】专家数量何为最优？面向混合专家模型的语义专业化优化研究

自进化人工智能体的全面综述：连接基础模型与终身自主智能系统的新范式

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

A Model-Based Approach for Specifying Changes in Replications of Empirical Studies in Computer Science

Arxiv

0+阅读 · 2022年6月27日

Towards Blockchain-Based Secure Data Management for Remote Patient Monitoring

Arxiv

0+阅读 · 2022年6月26日

Visual Auditor: Interactive Visualization for Detection and Summarization of Model Biases

Arxiv

0+阅读 · 2022年6月25日

Trustworthy AI: From Principles to Practices

Arxiv

46+阅读 · 2021年10月4日

Advances and Challenges in Conversational Recommender Systems: A Survey

Arxiv

14+阅读 · 2021年1月23日

相关基金

二层平衡问题的适定性与算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

对称性破缺条件下耦合系统chimera态的特性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Corin介导的ANP活化在动脉粥样硬化形成及其炎症反应中的作用与机制

国家自然科学基金

0+阅读 · 2012年12月31日

Z-pinch内爆等离子体二维高温辐射磁流体动力学方程及其动力系统研究

国家自然科学基金

0+阅读 · 2009年12月31日

冲击与火作用下弹塑性梁的动力响应及LTB研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员