FLEA:从不可靠的培训数据中吸取不可靠培训数据 (FLEA: Provably Fair Multisource Learning from Unreliable Training Data) - 专知论文

会员服务 ·

0

Facebook AI Research · 训练数据 · 学成 · 可辨认的 · 学习器 ·

2022 年 2 月 22 日

FLEA: Provably Fair Multisource Learning from Unreliable Training Data

翻译：FLEA:从不可靠的培训数据中吸取不可靠培训数据

Eugenia Iofinova,Nikola Konstantinov,Christoph H. Lampert

from arxiv, 12(main body)+41(appendix) pages, 0+20 figures. Latest version includes experimental results on the Folktables dataset

Fairness-aware learning aims at constructing classifiers that not only make accurate predictions, but also do not discriminate against specific groups. It is a fast-growing area of machine learning with far-reaching societal impact. However, existing fair learning methods are vulnerable to accidental or malicious artifacts in the training data, which can cause them to unknowingly produce unfair classifiers. In this work we address the problem of fair learning from unreliable training data in the robust multisource setting, where the available training data comes from multiple sources, a fraction of which might not be representative of the true data distribution. We introduce FLEA, a filtering-based algorithm that allows the learning system to identify and suppress those data sources that would have a negative impact on fairness or accuracy if they were used for training. We show the effectiveness of our approach by a diverse range of experiments on multiple datasets. Additionally, we prove formally that - given enough data - FLEA protects the learner against corruptions as long as the fraction of affected data sources is less than half.

翻译：公平了解学习旨在构建不仅作出准确预测,而且不歧视特定群体的分类师,这是一个快速增长的机器学习领域,具有深远的社会影响。然而,现有的公平学习方法在培训数据中容易发生意外或恶意文物,从而导致他们不知情地产生不公平的分类师。在这项工作中,我们处理从强有力的多来源环境中不可靠的培训数据中公平学习的问题,因为现有培训数据来自多个来源,其中一小部分可能无法代表真实的数据分布。我们引入了基于过滤的算法,使学习系统能够识别和抑制那些如果用于培训会对公平性或准确性产生消极影响的数据源。我们通过多种数据集的多种实验展示了我们的方法的有效性。此外,我们正式证明,只要受影响的数据源的比例不到一半,只要有了足够的数据,FLEA就能保护学习者免受腐败。

0

相关内容

Facebook AI Research

Facebook AI Research

Facebook AI Research

哥伦比亚大学最新《机器学习》课程，Fall-B 2020 (Machine Learning)

专知会员服务

39+阅读 · 2020年11月3日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

51+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

复杂工况下基于数据挖掘的资源消耗会计分摊方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

光纤LSPR传感器检测多元重金属的指纹识别与波长调控机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

面向柔性高分子太阳能电池的氧化石墨烯界面材料

国家自然科学基金

0+阅读 · 2012年12月31日

基于金属有机骨架材料的催化发光传感阵列研究

国家自然科学基金

1+阅读 · 2012年12月31日

云计算环境下的可信服务组合及运行保障研究

国家自然科学基金

0+阅读 · 2012年12月31日

整合常见和罕见变异进行肺癌风险预测的统计方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

纳米微粒/聚合物复合超粒子的构筑及功能

国家自然科学基金

0+阅读 · 2011年12月31日

金属巯基配合物的阴离子识别传感研究

国家自然科学基金

0+阅读 · 2009年12月31日

Unscented卡尔曼滤波算法及其在通信中的应用

国家自然科学基金

0+阅读 · 2008年12月31日

球面学习理论研究

国家自然科学基金

1+阅读 · 2008年12月31日

ARCH: Efficient Adversarial Regularized Training with Caching

Arxiv

0+阅读 · 2022年4月20日

Quantity vs Quality: Investigating the Trade-Off between Sample Size and Label Reliability

Arxiv

0+阅读 · 2022年4月20日

Case-Aware Adversarial Training

Arxiv

0+阅读 · 2022年4月20日

Named Entity Recognition for Partially Annotated Datasets

Arxiv

0+阅读 · 2022年4月19日

Active Learning Helps Pretrained Models Learn the Intended Task

Arxiv

1+阅读 · 2022年4月18日

FairFed: Enabling Group Fairness in Federated Learning

Arxiv

0+阅读 · 2022年4月18日

Bayesian Deep Learning for Graphs

Arxiv

23+阅读 · 2022年2月24日

The Causal Learning of Retail Delinquency

Arxiv

15+阅读 · 2020年12月17日

Learning from Very Few Samples: A Survey

Arxiv

126+阅读 · 2020年9月6日

Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks

Arxiv

14+阅读 · 2019年8月8日

VIP会员

文章信息

相关主题

Facebook AI Research

相关VIP内容

哥伦比亚大学最新《机器学习》课程，Fall-B 2020 (Machine Learning)

专知会员服务

39+阅读 · 2020年11月3日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

163+阅读 · 2019年10月12日

开源书：PyTorch深度学习起步

开源书：PyTorch深度学习起步

专知会员服务

51+阅读 · 2019年10月11日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《为多域数字战场变革装甲力量》报告

《多域训练：利用开放标准将太空与网络域同陆、海、空域训练相整合》报告

面向城市战：欧美徒步作战新装备

《人工智能增强监视分析：利用跨网络、陆地、空中及海上领域的威胁向量实时建模》

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium9

中国图象图形学学会CSIG

0+阅读 · 2021年12月17日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

【代码资源】GAN | 七份最热GAN文章及代码分享（Github 1000+Stars）

专知

13+阅读 · 2018年6月24日

相关论文

ARCH: Efficient Adversarial Regularized Training with Caching

Arxiv

0+阅读 · 2022年4月20日

Quantity vs Quality: Investigating the Trade-Off between Sample Size and Label Reliability

Arxiv

0+阅读 · 2022年4月20日

Case-Aware Adversarial Training

Arxiv

0+阅读 · 2022年4月20日

Named Entity Recognition for Partially Annotated Datasets

Arxiv

0+阅读 · 2022年4月19日

Active Learning Helps Pretrained Models Learn the Intended Task

Arxiv

1+阅读 · 2022年4月18日

FairFed: Enabling Group Fairness in Federated Learning

Arxiv

0+阅读 · 2022年4月18日

Bayesian Deep Learning for Graphs

Arxiv

23+阅读 · 2022年2月24日

The Causal Learning of Retail Delinquency

Arxiv

15+阅读 · 2020年12月17日

Learning from Very Few Samples: A Survey

Arxiv

126+阅读 · 2020年9月6日

Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks

Arxiv

14+阅读 · 2019年8月8日

相关基金

复杂工况下基于数据挖掘的资源消耗会计分摊方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

光纤LSPR传感器检测多元重金属的指纹识别与波长调控机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

面向柔性高分子太阳能电池的氧化石墨烯界面材料

国家自然科学基金

0+阅读 · 2012年12月31日

基于金属有机骨架材料的催化发光传感阵列研究

国家自然科学基金

1+阅读 · 2012年12月31日

云计算环境下的可信服务组合及运行保障研究

国家自然科学基金

0+阅读 · 2012年12月31日

整合常见和罕见变异进行肺癌风险预测的统计方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

纳米微粒/聚合物复合超粒子的构筑及功能

国家自然科学基金

0+阅读 · 2011年12月31日

金属巯基配合物的阴离子识别传感研究

国家自然科学基金

0+阅读 · 2009年12月31日

Unscented卡尔曼滤波算法及其在通信中的应用

国家自然科学基金

0+阅读 · 2008年12月31日

球面学习理论研究

国家自然科学基金

1+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员