反对亚洲仇恨的语音检测任务:BERT中心,而数据中心研究是关键 (Speech Detection Task Against Asian Hate: BERT the Central, While Data-Centric Studies the Crucial) - 专知论文

会员服务 ·

0

BERT · 数据集 · Performer · 讲稿 · Performance ·

2022 年 8 月 21 日

Speech Detection Task Against Asian Hate: BERT the Central, While Data-Centric Studies the Crucial

翻译：反对亚洲仇恨的语音检测任务:BERT中心,而数据中心研究是关键

With the COVID-19 pandemic continuing, hatred against Asians is intensifying in countries outside Asia, especially among the Chinese. There is an urgent need to detect and prevent hate speech towards Asians effectively. In this work, we first create COVID-HATE-2022, an annotated dataset including 2,025 annotated tweets fetched in early February 2022, which are labeled based on specific criteria, and we present the comprehensive collection of scenarios of hate and non-hate tweets in the dataset. Second, we fine-tune the BERT model based on the relevant datasets and demonstrate several strategies related to the "cleaning" of the tweets. Third, we investigate the performance of advanced fine-tuning strategies with various model-centric and data-centric approaches, and we show that both strategies generally improve the performance, while data-centric ones outperform the others, and it demonstrates the feasibility and effectiveness of the data-centric approaches in the associated tasks.

翻译：随着COVID-19大流行的继续,亚洲以外的国家,特别是中国,对亚洲人的仇恨正在加剧。迫切需要有效地发现和防止针对亚洲人的仇恨言论。在这项工作中,我们首先创建了COVID-HATE-2022,这是一个附加说明的数据集,包括2022年2月初收到的2 025条附加说明的推文,这些推文贴上了具体标准的标签,我们在数据集中全面收集仇恨和非仇恨推文的情景。第二,我们根据相关数据集对BERT模型进行微调,并展示了与“清理”这些推文有关的若干战略。第三,我们用各种以模型为中心的和以数据为中心的方法调查高级微调战略的绩效,我们表明,这两种战略总体上都改善了绩效,而以数据为中心的推文则优于其他标准。它展示了相关任务中以数据为中心的方法的可行性和有效性。

0

相关内容

BERT

BERT全称Bidirectional Encoder Representations from Transformers，是预训练语言表示的方法，可以在大型文本语料库（如维基百科）上训练通用的“语言理解”模型，然后将该模型用于下游NLP任务，比如机器翻译、问答。

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

转录激活蛋白YLGat1介导氮饥饿与油脂合成偶联的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

稀土上转换发光复合体系用于疾病的早期检测研究

国家自然科学基金

0+阅读 · 2015年12月31日

表面等离子体共振增强的新型高效太阳能电池

国家自然科学基金

0+阅读 · 2013年12月31日

AlCrN陶瓷薄膜韧性机制的原位原子尺度研究

国家自然科学基金

0+阅读 · 2013年12月31日

层层组装石墨烯复合材料的可控合成及其在太阳能电池中的应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

RegIII信号通路与SOCS3甲基化协同调控胰腺炎症恶性转化的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

两个WD40转录因子对银杏类黄酮生物合成调控的研究

国家自然科学基金

0+阅读 · 2012年12月31日

量子点和稀土离子共敏化二氧化钛纳米管阵列太阳能电池的研究

国家自然科学基金

0+阅读 · 2012年12月31日

新型抗生素Bagremycins生物合成基因簇的鉴定与解析

国家自然科学基金

0+阅读 · 2012年12月31日

树枝状分子功能性液晶凝胶的制备、性质与应用研究

国家自然科学基金

0+阅读 · 2009年12月31日

Augmentor or Filter? Reconsider the Role of Pre-trained Language Model in Text Classification Augmentation

Arxiv

0+阅读 · 2022年10月6日

Human-AI Shared Control via Policy Dissection

Arxiv

0+阅读 · 2022年10月5日

IoU-Enhanced Attention for End-to-End Task Specific Object Detection

Arxiv

0+阅读 · 2022年10月5日

Provable Guarantees against Data Poisoning Using Self-Expansion and Compatibility

Provable Guarantees against Data Poisoning Using Self-Expansion and Compatibility

Arxiv

0+阅读 · 2022年10月4日

What is the Price of a Skill? Revealing the Complementary Value of Skills

Arxiv

0+阅读 · 2022年10月4日

A Reproducible and Realistic Evaluation of Partial Domain Adaptation Methods

Arxiv

0+阅读 · 2022年10月3日

Unsupervised Model Selection for Time-series Anomaly Detection

Arxiv

0+阅读 · 2022年10月3日

Hypothesis Engineering for Zero-Shot Hate Speech Detection

Arxiv

0+阅读 · 2022年10月3日

The Dynamic of Consensus in Deep Networks and the Identification of Noisy Labels

Arxiv

0+阅读 · 2022年10月2日

Handle Anywhere: A Mobile Robot Arm for Providing Bodily Support to Elderly Persons

Arxiv

0+阅读 · 2022年9月30日

VIP会员

文章信息

相关主题

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

167+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

大语言模型智能体强化学习：全景综述

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

【伯克利博士论文】从推理服务到训练：面向大规模 LLM 智能体的高效系统

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

【ICIG2021】Latest News & Announcements of the Plenary Talk2

【ICIG2021】Latest News & Announcements of the Plenary Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年11月2日

【ICIG2021】Latest News & Announcements of the Industry Talk1

【ICIG2021】Latest News & Announcements of the Industry Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年7月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Augmentor or Filter? Reconsider the Role of Pre-trained Language Model in Text Classification Augmentation

Arxiv

0+阅读 · 2022年10月6日

Human-AI Shared Control via Policy Dissection

Arxiv

0+阅读 · 2022年10月5日

IoU-Enhanced Attention for End-to-End Task Specific Object Detection

Arxiv

0+阅读 · 2022年10月5日

Provable Guarantees against Data Poisoning Using Self-Expansion and Compatibility

Provable Guarantees against Data Poisoning Using Self-Expansion and Compatibility

Arxiv

0+阅读 · 2022年10月4日

What is the Price of a Skill? Revealing the Complementary Value of Skills

Arxiv

0+阅读 · 2022年10月4日

A Reproducible and Realistic Evaluation of Partial Domain Adaptation Methods

Arxiv

0+阅读 · 2022年10月3日

Unsupervised Model Selection for Time-series Anomaly Detection

Arxiv

0+阅读 · 2022年10月3日

Hypothesis Engineering for Zero-Shot Hate Speech Detection

Arxiv

0+阅读 · 2022年10月3日

The Dynamic of Consensus in Deep Networks and the Identification of Noisy Labels

Arxiv

0+阅读 · 2022年10月2日

Handle Anywhere: A Mobile Robot Arm for Providing Bodily Support to Elderly Persons

Arxiv

0+阅读 · 2022年9月30日

相关基金

转录激活蛋白YLGat1介导氮饥饿与油脂合成偶联的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

稀土上转换发光复合体系用于疾病的早期检测研究

国家自然科学基金

0+阅读 · 2015年12月31日

表面等离子体共振增强的新型高效太阳能电池

国家自然科学基金

0+阅读 · 2013年12月31日

AlCrN陶瓷薄膜韧性机制的原位原子尺度研究

国家自然科学基金

0+阅读 · 2013年12月31日

层层组装石墨烯复合材料的可控合成及其在太阳能电池中的应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

RegIII信号通路与SOCS3甲基化协同调控胰腺炎症恶性转化的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

两个WD40转录因子对银杏类黄酮生物合成调控的研究

国家自然科学基金

0+阅读 · 2012年12月31日

量子点和稀土离子共敏化二氧化钛纳米管阵列太阳能电池的研究

国家自然科学基金

0+阅读 · 2012年12月31日

新型抗生素Bagremycins生物合成基因簇的鉴定与解析

国家自然科学基金

0+阅读 · 2012年12月31日

树枝状分子功能性液晶凝胶的制备、性质与应用研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员