关于带有数据处理预处理的多重敏感属性的公平机器学习软件 (Fairer Machine Learning Software on Multiple Sensitive Attributes With Data Preprocessing) - 专知论文

会员服务 ·

0

Learning · Machine Learning · Facebook AI Research · 数据预处理 · Attention ·

2022 年 6 月 21 日

Fairer Machine Learning Software on Multiple Sensitive Attributes With Data Preprocessing

翻译：关于带有数据处理预处理的多重敏感属性的公平机器学习软件

Zhe Yu,Joymallya Chakraborty,Tim Menzies

from arxiv, 10 pages

This research seeks to benefit the software engineering society by providing a simple yet effective approach to improve fairness of machine learning software on data with multiple sensitive attributes. Machine learning fairness has attracted increasing attention since machine learning software is increasingly used for high-stakes and high-risk decisions. Amongst all the fairness notations, this work specifically targets "equalized odds". Equalized odds requires that members of every demographic group do not receive disparate mistreatment. It is one of the most widely accepted fairness notations given its advantage in always allowing perfect classifiers. Most existing solutions for machine learning fairness do not directly target equalized odds and only affect one sensitive attribute (e.g. sex) at a time. To overcome this shortage, we analyzed the condition of equalized odds and hypothesize that balancing the class distribution of training data across every demographic group will improve equalized odds of the learned model. On four real-world datasets (two of which have multiple sensitive attributes) and three synthetic datasets, our empirical results show that, at low computational overhead, the proposed preprocessing algorithm FairBalance can significantly improve equalized odds without much, if any damage to the prediction performance. FairBalance also outperforms existing state-of-the-art approaches in terms of equalized odds. To facilitate reuse, reproduction, and validation of this work, our scripts and data are available at https://github.com/hil-se/FairBalance under an open-source Apache license (v2.0).

翻译：这项研究旨在通过提供简单而有效的方法,提高机器学习软件对具有多种敏感属性的数据的公平性,从而让软件工程社会受益。由于机器学习软件越来越多地用于高考量和高风险决策,机器学习的公平性引起了越来越多的关注。在所有公平性评分中,这项工作具体针对的是“公平性差” 。公平性差要求每个人口群体的成员不受到不同的虐待。这是最广泛接受的公平性评分之一,因为它总能让完美的分类者受益。大多数现有的机器学习公平性办法并不直接针对均等机会,而且只影响到一个敏感属性(如性别)。为了克服这一短缺,我们分析了平衡每个人口群体培训数据班级分布的均等性差数和虚伪性状况,这将改善所学模型的均等性差数。关于四个真实世界数据集(其中两个具有多种敏感属性)和三个合成数据集,我们的实证结果显示,在低计算性间接费用下,拟议的预处理法公平性比差在一个时间点上可以大大改善均等性差,如果对预测性能造成任何损害的话。

0

相关内容

Learning

【图机器学习进展与趋势@ICML2022】Graph Machine Learning @ ICML 2022

【图机器学习进展与趋势@ICML2022】Graph Machine Learning @ ICML 2022

专知会员服务

40+阅读 · 2022年7月25日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

246+阅读 · 2019年10月21日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

LOC283683-NIPA1-BMPRII途径对胆固醇平衡和动脉粥样硬化的影响及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

机械合金化Nb-Al非平衡相中Nb3Al超导体析出机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Disabled-1可变性剪接在调控Reelin信号通路和神经元迁移中的分子研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

马尔可夫过程在Girsanov变换下的性质及其应用

国家自然科学基金

0+阅读 · 2012年12月31日

地上-地下的互作对入侵植物空心莲子草（Alternanthera philoxeroides）的影响及其响应机制

国家自然科学基金

0+阅读 · 2012年12月31日

TREM-1/DAP12/ NF-κB信号通路在6-姜烯酚抗动脉粥样硬化中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

Arisandilactone A 的不对称全合成

国家自然科学基金

0+阅读 · 2012年12月31日

Musclin基因在骨骼肌表达的转录调控机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

Modeling Extremal Streamflow using Deep Learning Approximations and a Flexible Spatial Process

Arxiv

0+阅读 · 2022年8月5日

Dynamic Adaptive and Adversarial Graph Convolutional Network for Traffic Forecasting

Arxiv

0+阅读 · 2022年8月5日

Fast Kernel Density Estimation with Density Matrices and Random Fourier Features

Arxiv

0+阅读 · 2022年8月4日

MAGPIE: Machine Automated General Performance Improvement via Evolution of Software

Arxiv

0+阅读 · 2022年8月4日

Exploration of Parameter Spaces Assisted by Machine Learning

Exploration of Parameter Spaces Assisted by Machine Learning

Arxiv

0+阅读 · 2022年8月4日

Bayesian Deep Learning for Graphs

Arxiv

23+阅读 · 2022年2月24日

Challenges of Artificial Intelligence -- From Machine Learning and Computer Vision to Emotional Intelligence

Arxiv

19+阅读 · 2022年1月5日

Automated Graph Machine Learning: Approaches, Libraries and Directions

Arxiv

20+阅读 · 2022年1月4日

Heterogeneous Network Representation Learning: A Unified Framework with Survey and Benchmark

Heterogeneous Network Representation Learning: A Unified Framework with Survey and Benchmark

Arxiv

19+阅读 · 2020年12月17日

Privacy and Robustness in Federated Learning: Attacks and Defenses

Arxiv

35+阅读 · 2020年12月7日

VIP会员

文章信息

相关主题

Machine Learning

Facebook AI Research

数据预处理

相关VIP内容

【图机器学习进展与趋势@ICML2022】Graph Machine Learning @ ICML 2022

【图机器学习进展与趋势@ICML2022】Graph Machine Learning @ ICML 2022

专知会员服务

40+阅读 · 2022年7月25日

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

【机器学习基础最新版】（Mathematics for Machine Learning），417页pdf

专知会员服务

246+阅读 · 2019年10月21日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】基础模型训练中网络规模数据的负责任与高效使用

《俄乌战争背景下俄罗斯的战略性海军分析（2022-2025年）》最新100页报告

人工智能时代背景下的未来海战

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Modeling Extremal Streamflow using Deep Learning Approximations and a Flexible Spatial Process

Arxiv

0+阅读 · 2022年8月5日

Dynamic Adaptive and Adversarial Graph Convolutional Network for Traffic Forecasting

Arxiv

0+阅读 · 2022年8月5日

Fast Kernel Density Estimation with Density Matrices and Random Fourier Features

Arxiv

0+阅读 · 2022年8月4日

MAGPIE: Machine Automated General Performance Improvement via Evolution of Software

Arxiv

0+阅读 · 2022年8月4日

Exploration of Parameter Spaces Assisted by Machine Learning

Exploration of Parameter Spaces Assisted by Machine Learning

Arxiv

0+阅读 · 2022年8月4日

Bayesian Deep Learning for Graphs

Arxiv

23+阅读 · 2022年2月24日

Challenges of Artificial Intelligence -- From Machine Learning and Computer Vision to Emotional Intelligence

Arxiv

19+阅读 · 2022年1月5日

Automated Graph Machine Learning: Approaches, Libraries and Directions

Arxiv

20+阅读 · 2022年1月4日

Heterogeneous Network Representation Learning: A Unified Framework with Survey and Benchmark

Heterogeneous Network Representation Learning: A Unified Framework with Survey and Benchmark

Arxiv

19+阅读 · 2020年12月17日

Privacy and Robustness in Federated Learning: Attacks and Defenses

Arxiv

35+阅读 · 2020年12月7日

相关基金

LOC283683-NIPA1-BMPRII途径对胆固醇平衡和动脉粥样硬化的影响及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

机械合金化Nb-Al非平衡相中Nb3Al超导体析出机理研究

国家自然科学基金

0+阅读 · 2013年12月31日

Disabled-1可变性剪接在调控Reelin信号通路和神经元迁移中的分子研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于SURE/PURE准则的图像盲反卷积算法研究

国家自然科学基金

3+阅读 · 2013年12月31日

马尔可夫过程在Girsanov变换下的性质及其应用

国家自然科学基金

0+阅读 · 2012年12月31日

地上-地下的互作对入侵植物空心莲子草（Alternanthera philoxeroides）的影响及其响应机制

国家自然科学基金

0+阅读 · 2012年12月31日

TREM-1/DAP12/ NF-κB信号通路在6-姜烯酚抗动脉粥样硬化中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

Arisandilactone A 的不对称全合成

国家自然科学基金

0+阅读 · 2012年12月31日

Musclin基因在骨骼肌表达的转录调控机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员