KNOW How to Make Up Your Mind! Adversarially Detecting and Alleviating Inconsistencies in Natural Language Explanations

While recent works have been considerably improving the quality of the natural language explanations (NLEs) generated by a model to justify its predictions, there is very limited research in detecting and alleviating inconsistencies among generated NLEs. In this work, we leverage external knowledge bases to significantly improve on an existing adversarial attack for detecting inconsistent NLEs. We apply our attack to high-performing NLE models and show that models with higher NLE quality do not necessarily generate fewer inconsistencies. Moreover, we propose an off-the-shelf mitigation method to alleviate inconsistencies by grounding the model into external background knowledge. Our method decreases the inconsistencies of previous high-performing NLE models as detected by our attack.

翻译：暂无翻译

相关内容

NLE

关注 200

自然语言工程（Natural Language Engineering）满足了自动语言处理各个领域的专业人员和研究人员的需求，无论是从理论还是语料库语言学、翻译、词典编纂、计算机科学还是工程学的角度。其目的是在传统的计算语言学研究和实际应用之间架起一座桥梁。除了出版关于广泛主题的原创研究文章——从文本分析、机器翻译、信息检索、语音处理和生成到集成系统和多模态接口——它还出版关于特定自然语言处理方法、任务或应用程序的特刊。官网地址：http://dblp.uni-trier.de/db/journals/nle/

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日