古中国单词分割和部分语音短语使用疏漏监管</s> (Ancient Chinese Word Segmentation and Part-of-Speech Tagging Using Distant Supervision)

Ancient Chinese word segmentation (WSG) and part-of-speech tagging (POS) are important to study ancient Chinese, but the amount of ancient Chinese WSG and POS tagging data is still rare. In this paper, we propose a novel augmentation method of ancient Chinese WSG and POS tagging data using distant supervision over parallel corpus. However, there are still mislabeled and unlabeled ancient Chinese words inevitably in distant supervision. To address this problem, we take advantage of the memorization effects of deep neural networks and a small amount of annotated data to get a model with much knowledge and a little noise, and then we use this model to relabel the ancient Chinese sentences in parallel corpus. Experiments show that the model trained over the relabeled data outperforms the model trained over the data generated from distant supervision and the annotated data. Our code is available at https://github.com/farlit/ACDS.

翻译：中国古代文字分割和部分语音标记对于研究古中国十分重要,但古中国WSG和POS标记数据的数量仍然很少。在本文中,我们提议采用新颖的增强方法,利用远处的平行保护系统对古中国WSG和POS数据进行标记,然而,在远处的监视下,仍然有误标和未贴标签的古代中文词句。为了解决这一问题,我们利用深层神经网络的记忆化效应和少量附加说明的数据来获得一个知识丰富、噪音小的模型,然后我们利用这个模型将古代中国句子重新标为平行体。实验显示,经过重新标签数据培训的模型比经过远程监督和附加说明数据培训的模型要强。我们的代码可在https://github.com/farlit/ACDS上查阅。</s>

相关内容

词性标注

关注 389

词性（part-of-speech）是词汇基本的语法属性，通常也称为词类。词性标注就是在给定句子中判定每个词的语法范畴，确定其词性并加以标注的过程，是中文信息处理面临的重要基础性问题。在语料库语言学中，词性标注（POS标注或PoS标注或POST），也称为语法标注，是将文本（语料库）中的单词标注为与特定词性相对应的过程，[1] 基于其定义和上下文。

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日