性别WWs:通过跨语言的语义化专业,在社交媒体中发现中国性别主义,从而实现域-警示单词嵌入 (SexWEs: Domain-Aware Word Embeddings via Cross-lingual Semantic Specialisation for Chinese Sexism Detection in Social Media)

The goal of sexism detection is to mitigate negative online content targeting certain gender groups of people. However, the limited availability of labeled sexism-related datasets makes it problematic to identify online sexism for low-resource languages. In this paper, we address the task of automatic sexism detection in social media for one low-resource language -- Chinese. Rather than collecting new sexism data or building cross-lingual transfer learning models, we develop a cross-lingual domain-aware semantic specialisation system in order to make the most of existing data. Semantic specialisation is a technique for retrofitting pre-trained distributional word vectors by integrating external linguistic knowledge (such as lexico-semantic relations) into the specialised feature space. To do this, we leverage semantic resources for sexism from a high-resource language (English) to specialise pre-trained word vectors in the target language (Chinese) to inject domain knowledge. We demonstrate the benefit of our sexist word embeddings (SexWEs) specialised by our framework via intrinsic evaluation of word similarity and extrinsic evaluation of sexism detection. Compared with other specialisation approaches and Chinese baseline word vectors, our SexWEs shows an average score improvement of 0.033 and 0.064 in both intrinsic and extrinsic evaluations, respectively. The ablative results and visualisation of SexWEs also prove the effectiveness of our framework on retrofitting word vectors in low-resource languages. Our code and sexism-related word vectors will be publicly available.

翻译：性别主义检测的目标是减少针对某些性别群体的负面在线内容。然而,由于贴标签的性别主义相关数据集数量有限,因此很难确定低资源语言的在线性别主义。在本文件中,我们处理的是社交媒体为一种低资源语言 -- -- 中文 -- -- 自动发现性别主义的任务。我们不是收集新的性别主义数据,或建立跨语言传输学习模式,而是开发一个跨语言域认知的静默专用系统,以便充分利用现有数据。语义专业化是一种技术,通过将外部语言知识(如词汇和语义关系)纳入特殊功能空间来改造预先培训的分发词矢量。为此,我们利用高资源语言(英语)在社交媒体中自动发现性别主义性别主义。我们不是收集新的性别主义数据,也不是建立跨语言(中文)预先培训的文字矢量转移学习模式,而是开发一种交叉语言嵌入低现有数据(SexWES)的好处。我们框架通过内在的词义相似性语言和极端语言矢量评估来改造预先定义传播语言的传播方式。在Syalimalalimalalimalalisalizalizalalalizalizalalizalization 和中国Salalalalalalalalalalismal devidustration 20 对比中, 以及中国Syalalalalalalalalalalalalizalalalalalalalalalalal 20 将展示一种其他特殊语言和Syalalalalalalalalalal 20 方法和自我化的词系的词系的自我标,在中国化的自我化的自我化和自我化和性代谢。 20 20 20 。我们性代谢性言,我们的性别系的性别系的词和SYalalalalalalalalalalalalalalalalalalalalalbalalalalalalalalalalalalalalalbalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalalbalalalalalalalalbalbalbalalalalalalalal