中国本土语系错误校正 (Linguistic Rules-Based Corpus Generation for Native Chinese Grammatical Error Correction) - 专知论文

会员服务 ·

0

Performer · MoDELS · Extensibility · Excel · 训练数据 ·

2022 年 10 月 19 日

Linguistic Rules-Based Corpus Generation for Native Chinese Grammatical Error Correction

翻译：中国本土语系错误校正

Shirong Ma,Yinghui Li,Rongyi Sun,Qingyu Zhou,Shulin Huang,Ding Zhang,Li Yangning,Ruiyang Liu,Zhongli Li,Yunbo Cao,Haitao Zheng,Ying Shen

from arxiv, Long paper, accepted at the Findings of EMNLP 2022

Chinese Grammatical Error Correction (CGEC) is both a challenging NLP task and a common application in human daily life. Recently, many data-driven approaches are proposed for the development of CGEC research. However, there are two major limitations in the CGEC field: First, the lack of high-quality annotated training corpora prevents the performance of existing CGEC models from being significantly improved. Second, the grammatical errors in widely used test sets are not made by native Chinese speakers, resulting in a significant gap between the CGEC models and the real application. In this paper, we propose a linguistic rules-based approach to construct large-scale CGEC training corpora with automatically generated grammatical errors. Additionally, we present a challenging CGEC benchmark derived entirely from errors made by native Chinese speakers in real-world scenarios. Extensive experiments and detailed analyses not only demonstrate that the training data constructed by our method effectively improves the performance of CGEC models, but also reflect that our benchmark is an excellent resource for further development of the CGEC field.

翻译：中国典型错误校正(CGEC)既是一项艰巨的任务,也是人类日常生活中的一种常见应用。最近,提出了许多以数据为驱动的办法来发展个体分类研究。然而,在个体分类领域存在两大限制:第一,缺乏高质量的附加说明的培训公司,使得现有个体分类模型的性能无法大大改进。第二,广用测试组中的语法错误不是由本地中文演讲人造成的,造成个体分类模型与实际应用之间的巨大差距。在本文件中,我们提出了一种基于语言的基于规则的方法,用自动生成的语法错误来构建大型个体分类公司的培训。此外,我们提出了一个具有挑战性的个体分类中心基准,完全源于本地中文演讲人在现实世界情景中所犯的错误。广泛的实验和详细分析不仅表明我们的方法构建的培训数据有效地改进了个体分类模型的性能,还反映出我们的基准是进一步开发个体分类领域的最佳资源。

0

相关内容

Performer

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

高磷血症致胆固醇敏感器SCAP功能失调促进动脉粥样硬化的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

雷公藤甲素诱导急性早幼粒白血病细胞凋亡及自噬的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

领域驱动空间co-location模式挖掘技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

胆固醇转运子ABCA1调控CD4+T细胞免疫应答抑制动脉粥样硬化的新机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于混合网格向量值细分曲面可计算光滑性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Diversin介导非小细胞肺癌长春瑞滨耐药的分子机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

新城疫病毒感染鸡树突状细胞抑制T淋巴细胞增殖作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Levy扩散过程与非局部偏微分方程

国家自然科学基金

1+阅读 · 2012年12月31日

CliMedBERT: A Pre-trained Language Model for Climate and Health-related Text

Arxiv

0+阅读 · 2022年12月1日

Leveraging Large-scale Multimedia Datasets to Refine Content Moderation Models

Leveraging Large-scale Multimedia Datasets to Refine Content Moderation Models

Arxiv

0+阅读 · 2022年12月1日

Language Model Pre-training on True Negatives

Arxiv

0+阅读 · 2022年12月1日

How to Train an Accurate and Efficient Object Detection Model on Any Dataset

Arxiv

0+阅读 · 2022年11月30日

BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch?

Arxiv

0+阅读 · 2022年11月30日

Action-GPT: Leveraging Large-scale Language Models for Improved and Generalized Zero Shot Action Generation

Arxiv

0+阅读 · 2022年11月30日

Automated Generating Natural Language Requirements based on Domain Ontology

Arxiv

0+阅读 · 2022年11月30日

An Experiment Design Paradigm using Joint Feature Selection and Task Optimization

Arxiv

0+阅读 · 2022年11月29日

Adversarial Mutual Information for Text Generation

Adversarial Mutual Information for Text Generation

Arxiv

13+阅读 · 2020年6月30日

Extreme Language Model Compression with Optimal Subwords and Shared Projections

Extreme Language Model Compression with Optimal Subwords and Shared Projections

Arxiv

18+阅读 · 2019年9月25日

VIP会员

文章信息

相关主题

相关VIP内容

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

小规模训练指南：打造世界级大语言模型的关键方法

无人机编队飞行：复杂环境中作战的策略、挑战与应用

大模型APP，AI时代第一个爆款

从数据中心视角出发的高效大语言模型训练综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

IEEE TII Call For Papers

IEEE TII Call For Papers

CCF多媒体专委会

3+阅读 · 2022年3月24日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

CliMedBERT: A Pre-trained Language Model for Climate and Health-related Text

Arxiv

0+阅读 · 2022年12月1日

Leveraging Large-scale Multimedia Datasets to Refine Content Moderation Models

Leveraging Large-scale Multimedia Datasets to Refine Content Moderation Models

Arxiv

0+阅读 · 2022年12月1日

Language Model Pre-training on True Negatives

Arxiv

0+阅读 · 2022年12月1日

How to Train an Accurate and Efficient Object Detection Model on Any Dataset

Arxiv

0+阅读 · 2022年11月30日

BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch?

Arxiv

0+阅读 · 2022年11月30日

Action-GPT: Leveraging Large-scale Language Models for Improved and Generalized Zero Shot Action Generation

Arxiv

0+阅读 · 2022年11月30日

Automated Generating Natural Language Requirements based on Domain Ontology

Arxiv

0+阅读 · 2022年11月30日

An Experiment Design Paradigm using Joint Feature Selection and Task Optimization

Arxiv

0+阅读 · 2022年11月29日

Adversarial Mutual Information for Text Generation

Adversarial Mutual Information for Text Generation

Arxiv

13+阅读 · 2020年6月30日

Extreme Language Model Compression with Optimal Subwords and Shared Projections

Extreme Language Model Compression with Optimal Subwords and Shared Projections

Arxiv

18+阅读 · 2019年9月25日

相关基金

高磷血症致胆固醇敏感器SCAP功能失调促进动脉粥样硬化的分子机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

雷公藤甲素诱导急性早幼粒白血病细胞凋亡及自噬的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

领域驱动空间co-location模式挖掘技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

胆固醇转运子ABCA1调控CD4+T细胞免疫应答抑制动脉粥样硬化的新机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于混合网格向量值细分曲面可计算光滑性研究

国家自然科学基金

0+阅读 · 2013年12月31日

Calderon问题和边界刚性问题

国家自然科学基金

0+阅读 · 2013年12月31日

Diversin介导非小细胞肺癌长春瑞滨耐药的分子机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

新城疫病毒感染鸡树突状细胞抑制T淋巴细胞增殖作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Levy扩散过程与非局部偏微分方程

国家自然科学基金

1+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员