评估自然语言处理（NLP）技术的多样性、公平性和包容性：印度语言的案例研究 (Evaluating the Diversity, Equity and Inclusion of NLP Technology: A Case Study for Indian Languages) - 专知论文

会员服务 ·

0

公平性 · 多样性 · NLP · 优化资源分配 · 语言处理 ·

2023 年 4 月 12 日

Evaluating the Diversity, Equity and Inclusion of NLP Technology: A Case Study for Indian Languages

翻译：评估自然语言处理（NLP）技术的多样性、公平性和包容性：印度语言的案例研究

Simran Khanuja,Sebastian Ruder,Partha Talukdar

from arxiv, Accepted to EACL Findings, 2023

In order for NLP technology to be widely applicable, fair, and useful, it needs to serve a diverse set of speakers across the world's languages, be equitable, i.e., not unduly biased towards any particular language, and be inclusive of all users, particularly in low-resource settings where compute constraints are common. In this paper, we propose an evaluation paradigm that assesses NLP technologies across all three dimensions. While diversity and inclusion have received attention in recent literature, equity is currently unexplored. We propose to address this gap using the Gini coefficient, a well-established metric used for estimating societal wealth inequality. Using our paradigm, we highlight the distressed state of current technologies for Indian (IN) languages (a linguistically large and diverse set, with a varied speaker population), across all three dimensions. To improve upon these metrics, we demonstrate the importance of region-specific choices in model building and dataset creation, and more importantly, propose a novel, generalisable approach to optimal resource allocation during fine-tuning. Finally, we discuss steps to mitigate these biases and encourage the community to employ multi-faceted evaluation when building linguistically diverse and equitable technologies.

翻译：为了使 NLP 技术具有广泛的适用性、公平性和实用性，它需要为全球语言中的多样化的使用者提供服务，具有公平性，即不偏袒任何特定的语言，并包容所有用户，特别是在计算受限制的低资源环境下。本文提出了一种评估范式，以对所有三个维度的 NLP 技术进行评估。尽管多样性和包容性在最近的文献中受到了关注，但公平性目前尚未得到探索。我们建议使用洪武系数，这是一种用于估算社会财富不平等的公认度量标准。使用我们的范式，我们突出了目前适用于印度语言 (IN) 的技术在所有三个维度上的贫困状态（印度语言是一种语言庞大多样，使用者群体多样化的语言）。为了改善这些指标，我们展示了地区特定选择在模型构建和数据集创建方面的重要性，更重要的是，我们提出了一种新颖的、可推广的方法来进行优化资源分配时的微调。最后，我们讨论了缓解这些偏见的措施，并鼓励社区在构建语言多样和公平的技术时采用多方面的评估。

0

相关内容

公平性

【2023新书】生成式AI和ChatGPT的兴起:了解生成式AI和ChatGPT如何改变和重塑商业世界，269页pdf

【2023新书】生成式AI和ChatGPT的兴起:了解生成式AI和ChatGPT如何改变和重塑商业世界，269页pdf

专知会员服务

111+阅读 · 2023年5月26日

【SIGMOD教程】高效数据标签的众包实践:聚合、增量重标签和定价，附180页slides

【SIGMOD教程】高效数据标签的众包实践:聚合、增量重标签和定价，附180页slides

专知会员服务

11+阅读 · 2022年10月20日

【元宇宙】“The State Of The Metaverse”26页报告

【元宇宙】“The State Of The Metaverse”26页报告

专知会员服务

45+阅读 · 2022年5月25日

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

139+阅读 · 2020年7月10日

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

专知会员服务

65+阅读 · 2020年5月12日

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

专知会员服务

43+阅读 · 2020年4月22日

【CMU-TACL2020】低资源跨语言实体链接，Low-resource Crosslingual EntityLinking

专知会员服务

17+阅读 · 2020年3月29日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

ACL 2022 | 基于Prompt的自动去偏：有效减轻预训练语言模型中的偏见

ACL 2022 | 基于Prompt的自动去偏：有效减轻预训练语言模型中的偏见

PaperWeekly

0+阅读 · 2022年7月14日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

【KDD2020-Tutorial】深度学习异常检测，180页ppt

【KDD2020-Tutorial】深度学习异常检测，180页ppt

专知

49+阅读 · 2020年8月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

NLP 2018 Highlights：2018自然语言处理技术亮点汇总

NLP 2018 Highlights：2018自然语言处理技术亮点汇总

AINLP

10+阅读 · 2019年2月9日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

【论文推荐】最新六篇自动问答相关论文—无监督迁移学习、综述、生成式问答、QDEE、可扩展文档理解

【论文推荐】最新六篇自动问答相关论文—无监督迁移学习、综述、生成式问答、QDEE、可扩展文档理解

专知

12+阅读 · 2018年5月9日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

自然语言处理 (NLP)资源大全

自然语言处理 (NLP)资源大全

机械鸡

35+阅读 · 2017年9月17日

非主从式混合云存储系统伸缩性管理研究

国家自然科学基金

1+阅读 · 2015年12月31日

氧化石墨烯对植物病原真菌的杀菌机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

黔西南高砷煤矿区环境重金属污染特征

国家自然科学基金

0+阅读 · 2014年12月31日

面向企业的商品评论代表性意见提取策略研究

国家自然科学基金

0+阅读 · 2013年12月31日

IPD项目交付模式下风险分担机制与理论模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

新兴市场国家IFRS制定过程中的博弈及经济后果研究

国家自然科学基金

1+阅读 · 2012年12月31日

白云鄂博矿区苔藓植物多样性及其对稀土元素的富集特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向属性的CPN建模及On the Fly辅助的测试生成方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于Bregman距离的一致性风险测度及其应用

国家自然科学基金

0+阅读 · 2011年12月31日

Reality-based Interaction用户界面模型和评估方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

The Magic of IF: Investigating Causal Reasoning Abilities in Large Language Models of Code

Arxiv

0+阅读 · 2023年5月30日

The Utility of Large Language Models and Generative AI for Education Research

Arxiv

0+阅读 · 2023年5月29日

The Leximin Approach for a Sequence of Collective Decisions

Arxiv

0+阅读 · 2023年5月29日

Predicting Survey Response with Quotation-based Modeling: A Case Study on Favorability towards the United States

Arxiv

0+阅读 · 2023年5月27日

Improving Stability in Decision Tree Models

Arxiv

0+阅读 · 2023年5月26日

Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person

Arxiv

0+阅读 · 2023年5月26日

Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Arxiv

0+阅读 · 2023年5月26日

Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review

Arxiv

0+阅读 · 2023年5月26日

Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond

Arxiv

21+阅读 · 2021年9月2日

Recent Advances in Deep Learning-based Dialogue Systems

Arxiv

18+阅读 · 2021年5月10日

VIP会员

文章信息

相关主题

优化资源分配

相关VIP内容

【2023新书】生成式AI和ChatGPT的兴起:了解生成式AI和ChatGPT如何改变和重塑商业世界，269页pdf

【2023新书】生成式AI和ChatGPT的兴起:了解生成式AI和ChatGPT如何改变和重塑商业世界，269页pdf

专知会员服务

111+阅读 · 2023年5月26日

【SIGMOD教程】高效数据标签的众包实践:聚合、增量重标签和定价，附180页slides

【SIGMOD教程】高效数据标签的众包实践:聚合、增量重标签和定价，附180页slides

专知会员服务

11+阅读 · 2022年10月20日

【元宇宙】“The State Of The Metaverse”26页报告

【元宇宙】“The State Of The Metaverse”26页报告

专知会员服务

45+阅读 · 2022年5月25日

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

最新《自然语言处理迁移学习》综述论文，A Survey on Transfer Learning in Natural Language Processing

专知会员服务

139+阅读 · 2020年7月10日

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

【视频描述综述论文】Video Description: A Survey of Methods, Datasets, and Evaluation Metrics

专知会员服务

65+阅读 · 2020年5月12日

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

【NLP模型压缩方法综述】《A Survey of Methods for Model Compression in NLP》by Madison May

专知会员服务

43+阅读 · 2020年4月22日

【CMU-TACL2020】低资源跨语言实体链接，Low-resource Crosslingual EntityLinking

专知会员服务

17+阅读 · 2020年3月29日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

热门VIP内容

开通专知VIP会员享更多权益服务

《美陆军特种作战条令》最新102页

《洛克希德SR-71“黑鸟”侦察机动力系统》21页slides

美空军作战实验室通过人工智能和指挥控制技术创新推进杀伤链

《指挥控制能力分析方法论》最新报告

相关资讯

ACL 2022 | 基于Prompt的自动去偏：有效减轻预训练语言模型中的偏见

ACL 2022 | 基于Prompt的自动去偏：有效减轻预训练语言模型中的偏见

PaperWeekly

0+阅读 · 2022年7月14日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

【KDD2020-Tutorial】深度学习异常检测，180页ppt

【KDD2020-Tutorial】深度学习异常检测，180页ppt

专知

49+阅读 · 2020年8月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

NLP 2018 Highlights：2018自然语言处理技术亮点汇总

NLP 2018 Highlights：2018自然语言处理技术亮点汇总

AINLP

10+阅读 · 2019年2月9日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

【论文推荐】最新六篇自动问答相关论文—无监督迁移学习、综述、生成式问答、QDEE、可扩展文档理解

【论文推荐】最新六篇自动问答相关论文—无监督迁移学习、综述、生成式问答、QDEE、可扩展文档理解

专知

12+阅读 · 2018年5月9日

【推荐】自然语言处理（NLP）指南

【推荐】自然语言处理（NLP）指南

机器学习研究会

35+阅读 · 2017年11月17日

自然语言处理 (NLP)资源大全

自然语言处理 (NLP)资源大全

机械鸡

35+阅读 · 2017年9月17日

相关论文

The Magic of IF: Investigating Causal Reasoning Abilities in Large Language Models of Code

Arxiv

0+阅读 · 2023年5月30日

The Utility of Large Language Models and Generative AI for Education Research

Arxiv

0+阅读 · 2023年5月29日

The Leximin Approach for a Sequence of Collective Decisions

Arxiv

0+阅读 · 2023年5月29日

Predicting Survey Response with Quotation-based Modeling: A Case Study on Favorability towards the United States

Arxiv

0+阅读 · 2023年5月27日

Improving Stability in Decision Tree Models

Arxiv

0+阅读 · 2023年5月26日

Evaluating OpenAI's Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person

Arxiv

0+阅读 · 2023年5月26日

Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

Arxiv

0+阅读 · 2023年5月26日

Towards a Common Understanding of Contributing Factors for Cross-Lingual Transfer in Multilingual Language Models: A Review

Arxiv

0+阅读 · 2023年5月26日

Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond

Arxiv

21+阅读 · 2021年9月2日

Recent Advances in Deep Learning-based Dialogue Systems

Arxiv

18+阅读 · 2021年5月10日

相关基金

非主从式混合云存储系统伸缩性管理研究

国家自然科学基金

1+阅读 · 2015年12月31日

氧化石墨烯对植物病原真菌的杀菌机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

黔西南高砷煤矿区环境重金属污染特征

国家自然科学基金

0+阅读 · 2014年12月31日

面向企业的商品评论代表性意见提取策略研究

国家自然科学基金

0+阅读 · 2013年12月31日

IPD项目交付模式下风险分担机制与理论模型研究

国家自然科学基金

0+阅读 · 2012年12月31日

新兴市场国家IFRS制定过程中的博弈及经济后果研究

国家自然科学基金

1+阅读 · 2012年12月31日

白云鄂博矿区苔藓植物多样性及其对稀土元素的富集特性研究

国家自然科学基金

0+阅读 · 2012年12月31日

面向属性的CPN建模及On the Fly辅助的测试生成方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

基于Bregman距离的一致性风险测度及其应用

国家自然科学基金

0+阅读 · 2011年12月31日

Reality-based Interaction用户界面模型和评估方法研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员