CLASP: 用于语义分析的鲜热交叉数据增强 (CLASP: Few-Shot Cross-Lingual Data Augmentation for Semantic Parsing) - 专知论文

会员服务 ·

0

语义分析 · MoDELS · 小样本学习 · 训练数据 · 数据增强 ·

2022 年 10 月 14 日

CLASP: Few-Shot Cross-Lingual Data Augmentation for Semantic Parsing

翻译：CLASP: 用于语义分析的鲜热交叉数据增强

Andy Rosenbaum,Saleh Soltan,Wael Hamza,Amir Saffari,Marco Damonte,Isabel Groves

from arxiv, Accepted to AACL-IJCNLP 2022: The 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, November 20-23, 2022. See https://www.aacl2022.org/

A bottleneck to developing Semantic Parsing (SP) models is the need for a large volume of human-labeled training data. Given the complexity and cost of human annotation for SP, labeled data is often scarce, particularly in multilingual settings. Large Language Models (LLMs) excel at SP given only a few examples, however LLMs are unsuitable for runtime systems which require low latency. In this work, we propose CLASP, a simple method to improve low-resource SP for moderate-sized models: we generate synthetic data from AlexaTM 20B to augment the training set for a model 40x smaller (500M parameters). We evaluate on two datasets in low-resource settings: English PIZZA, containing either 348 or 16 real examples, and mTOP cross-lingual zero-shot, where training data is available only in English, and the model must generalize to four new languages. On both datasets, we show significant improvements over strong baseline methods.

翻译：开发语义分解(SP)模型的一个瓶颈是需要大量的人类标签培训数据。鉴于人类对SP的批注的复杂性和成本,标签数据往往很少,特别是在多语种环境中。大型语言模型(LLMS)在SP中仅举几个例子,大语言模型(LLMS)优于SP,但是LLMS不适合运行时间系统,而运行时间系统需要低潜伏。在这项工作中,我们建议CLASP,这是改进中小型模型的低资源SP的一个简单方法:我们从AlexaTM 20B中生成合成数据,以扩大模型40x较小(500M参数)的培训集。我们评估了在低资源环境中的两个数据集:英文 PIZZA, 包含348或16个真实例子,以及 mTOP跨语言零弹,只有英语培训数据,而模型必须概括为四种新语言。在这两个数据集中,我们展示了强基线方法方面的显著改进。

0

相关内容

语义分析

语义分析的最终目的是理解句子表达的真实语义。但是，语义应该采用什么表示形式一直困扰着研究者们，至今这个问题也没有一个统一的答案。语义角色标注（semantic role labeling）是目前比较成熟的浅层语义分析技术。基于逻辑表达的语义分析也得到学术界的长期关注。

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

专知会员服务

21+阅读 · 2022年3月18日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

温肺化纤汤介导Wnt经典信号通路调控骨髓间充质干细胞向Ⅱ型肺泡细胞分化的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

miR-29b在Ang-II诱导肾小管上皮间充质转分化中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

低氧下白细胞介素-1β在肿瘤相关巨噬细胞介导肝癌上皮间质化中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

片仔癀调控microRNA抑制结肠癌上皮间质转化的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

p38信号通路调节对DC疫苗免疫治疗骨髓瘤的作用和机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Snai1/slug-miR30a反馈环路对肾小管上皮细胞间质转化的调控

国家自然科学基金

0+阅读 · 2012年12月31日

武汉“+1”#22478;市圈土地资源优化配置研究

国家自然科学基金

0+阅读 · 2011年12月31日

核外ATM蛋白在脂蛋白内吞过程中的作用及机制

国家自然科学基金

0+阅读 · 2009年12月31日

增强现实中多目标3D跟踪定位和WH-SIFT特征识别方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

Notch1信号通路在HBx致肝细胞恶性转化中的作用及机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

ClaSP -- Parameter-free Time Series Segmentation

Arxiv

0+阅读 · 2022年11月18日

Conffusion: Confidence Intervals for Diffusion Models

Arxiv

0+阅读 · 2022年11月17日

ConNER: Consistency Training for Cross-lingual Named Entity Recognition

Arxiv

0+阅读 · 2022年11月17日

Calibrated Interpretation: Confidence Estimation in Semantic Parsing

Arxiv

0+阅读 · 2022年11月16日

Streaming Joint Speech Recognition and Disfluency Detection

Arxiv

0+阅读 · 2022年11月16日

Adapting Pretrained Text-to-Text Models for Long Text Sequences

Arxiv

0+阅读 · 2022年11月16日

Interclass Prototype Relation for Few-Shot Segmentation

Arxiv

0+阅读 · 2022年11月16日

MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

Arxiv

0+阅读 · 2022年11月15日

Context-Matched Collage Generation for Underwater Invertebrate Detection

Arxiv

0+阅读 · 2022年11月15日

Deep Representation Learning for Domain Adaptation of Semantic Image Segmentation

Arxiv

10+阅读 · 2018年5月10日

VIP会员

文章信息

相关主题

小样本学习

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

专知会员服务

21+阅读 · 2022年3月18日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】以人为中心的强化学习

任务规划与地形分析：现代复杂环境作战导航体系

认知优势：人工智能在国家安全决策中的核心作用

大模型赋能的具身智能：决策与具身学习综述

相关资讯

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

征稿 | International Joint Conference on Knowledge Graphs (IJCKG)

开放知识图谱

2+阅读 · 2022年5月20日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

【论文推荐】最新6篇视觉问答（VQA）相关论文—目标推理、深度循环模型、可解释性、数据可视化、Triplet学习、基准

专知

15+阅读 · 2018年2月3日

相关论文

ClaSP -- Parameter-free Time Series Segmentation

Arxiv

0+阅读 · 2022年11月18日

Conffusion: Confidence Intervals for Diffusion Models

Arxiv

0+阅读 · 2022年11月17日

ConNER: Consistency Training for Cross-lingual Named Entity Recognition

Arxiv

0+阅读 · 2022年11月17日

Calibrated Interpretation: Confidence Estimation in Semantic Parsing

Arxiv

0+阅读 · 2022年11月16日

Streaming Joint Speech Recognition and Disfluency Detection

Arxiv

0+阅读 · 2022年11月16日

Adapting Pretrained Text-to-Text Models for Long Text Sequences

Arxiv

0+阅读 · 2022年11月16日

Interclass Prototype Relation for Few-Shot Segmentation

Arxiv

0+阅读 · 2022年11月16日

MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

Arxiv

0+阅读 · 2022年11月15日

Context-Matched Collage Generation for Underwater Invertebrate Detection

Arxiv

0+阅读 · 2022年11月15日

Deep Representation Learning for Domain Adaptation of Semantic Image Segmentation

Arxiv

10+阅读 · 2018年5月10日

相关基金

温肺化纤汤介导Wnt经典信号通路调控骨髓间充质干细胞向Ⅱ型肺泡细胞分化的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

miR-29b在Ang-II诱导肾小管上皮间充质转分化中的作用

国家自然科学基金

0+阅读 · 2013年12月31日

低氧下白细胞介素-1β在肿瘤相关巨噬细胞介导肝癌上皮间质化中的作用研究

国家自然科学基金

0+阅读 · 2013年12月31日

片仔癀调控microRNA抑制结肠癌上皮间质转化的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

p38信号通路调节对DC疫苗免疫治疗骨髓瘤的作用和机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

Snai1/slug-miR30a反馈环路对肾小管上皮细胞间质转化的调控

国家自然科学基金

0+阅读 · 2012年12月31日

武汉“+1”#22478;市圈土地资源优化配置研究

国家自然科学基金

0+阅读 · 2011年12月31日

核外ATM蛋白在脂蛋白内吞过程中的作用及机制

国家自然科学基金

0+阅读 · 2009年12月31日

增强现实中多目标3D跟踪定位和WH-SIFT特征识别方法研究

国家自然科学基金

0+阅读 · 2009年12月31日

Notch1信号通路在HBx致肝细胞恶性转化中的作用及机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员