跨语言情感检测 (Cross-lingual Emotion Detection)

Emotion detection can provide us with a window into understanding human behavior. Due to the complex dynamics of human emotions, however, constructing annotated datasets to train automated models can be expensive. Thus, we explore the efficacy of cross-lingual approaches that would use data from a source language to build models for emotion detection in a target language. We compare three approaches, namely: i) using inherently multilingual models; ii) translating training data into the target language; and iii) using an automatically tagged parallel corpus. In our study, we consider English as the source language with Arabic and Spanish as target languages. We study the effectiveness of different classification models such as BERT and SVMs trained with different features. Our BERT-based monolingual models that are trained on target language data surpass state-of-the-art (SOTA) by 4% and 5% absolute Jaccard score for Arabic and Spanish respectively. Next, we show that using cross-lingual approaches with English data alone, we can achieve more than 90% and 80% relative effectiveness of the Arabic and Spanish BERT models respectively. Lastly, we use LIME to analyze the challenges of training cross-lingual models for different language pairs

翻译：感官检测可以为我们提供理解人类行为的窗口。但是,由于人类情感的复杂动态,建造附加说明的数据集以培训自动化模型可能费用高昂。因此,我们探索了跨语言方法的功效,这些方法将使用源语言的数据来构建一种目标语言的情绪检测模型。我们比较了三种方法,即:一)使用固有的多语言模型;二)将培训数据转换成目标语言;三)使用自动标记的平行材料;我们的研究认为英语是源语言,阿拉伯语和西班牙语是目标语言。我们研究了不同分类模式的有效性,如BERT和受过不同特征培训的SVMs。我们基于BERT的单语模式,在目标语言数据方面接受培训,其语言数据比阿拉伯语和西班牙语分别高出4%和5%的绝对Jaccard分数。其次,我们显示,仅使用英语数据使用跨语言方法,我们就能分别实现阿拉伯语和西班牙语BERT模型90%和80%的相对有效性。最后,我们利用LME分析不同语言组合培训跨语言模型的挑战。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

20篇「ACL2020」最新论文抢先看！看自然语言处理2020在研究什么？

专知会员服务

97+阅读 · 2020年4月10日

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

【AAAI2020】Context-Transformer:上下文转换器:解决对象混淆的小样本检测，Context-Transformer: Tackling Object Confusion for Few-Shot Detection

专知会员服务

51+阅读 · 2020年3月17日