以遗忘的因果语言模型实现更好少点热和微调性能 (Towards Better Few-Shot and Finetuning Performance with Forgetful Causal Language Models) - 专知论文

会员服务 ·

0

Performer · 语言模型化 · 小样本学习 · 词元分析器 · MoDELS ·

2023 年 1 月 31 日

Towards Better Few-Shot and Finetuning Performance with Forgetful Causal Language Models

翻译：以遗忘的因果语言模型实现更好少点热和微调性能

Hao Liu,Xinyang Geng,Lisa Lee,Igor Mordatch,Sergey Levine,Sharan Narang,Pieter Abbeel

from arxiv, Added T-FCM and better FCM results

Large language models (LLM) trained using the next-token-prediction objective, such as GPT3 and PaLM, have revolutionized natural language processing in recent years by showing impressive zero-shot and few-shot capabilities across a wide range of tasks. In this work, we propose a simple technique that significantly boosts the performance of LLMs without adding computational cost. Our key observation is that, by performing the next token prediction task with randomly selected past tokens masked out, we can improve the quality of the learned representations for downstream language understanding tasks. We hypothesize that randomly masking past tokens prevents over-attending to recent tokens and encourages attention to tokens in the distant past. We find that our method, Forgetful Causal Masking (FCM), significantly improves both few-shot and finetuning performance of PaLM. We further consider a simple extension, T-FCM, which introduces bidirectional context to causal language model without altering the sequence order, and further improves finetuning performance.

翻译：使用下端口令目标(如GPT3和PALM)培训的大型语言模型(LLM)近年来通过在一系列任务中显示令人印象深刻的零射和几射能力,使自然语言处理发生了革命性的变化。在这项工作中,我们提出了一个简单的方法,在不增加计算成本的情况下大大提升了LLM的性能。我们的主要观察是,通过以随机选择的过去代号来完成下一个象征性的预测任务,我们可以提高为下游语言理解任务所学习的演示的质量。我们假设,随机遮盖过去代号防止过度使用最近的代号,并鼓励注意远古代代代代号。我们发现,我们的方法,即遗忘的Causal蒙码(FCM),大大改进了PALM的微光和微调性性能。我们进一步考虑一个简单的扩展,即T-FCM,在不改变序列顺序顺序的情况下引入因果关系语言模型的双向环境,并进一步改进性能。

0

相关内容

Performer

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

机器学习组合优化

机器学习组合优化

专知会员服务

110+阅读 · 2021年2月16日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

胡杨耐盐的表观基因组学研究

国家自然科学基金

0+阅读 · 2015年12月31日

茉莉酸调控水稻花器官发育分子调控网络的研究

国家自然科学基金

0+阅读 · 2014年12月31日

自噬对高脂膳食诱导的血管内皮细胞损伤的保护作用及分子机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

LPA-LPAR1信号通路调控放射性肺损伤的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Rbp1在特络细胞联合骨髓间充质干细胞治疗急性肺损伤中的调节作用和分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

混凝土细观损伤模拟与数值尺寸效应研究

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预NF-кB信号通路防治动脉粥样硬化的作用及机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

遍历哈密顿系统的谱理论

国家自然科学基金

0+阅读 · 2009年12月31日

C3G－SOD双靶点阻遏ROS多环节改善PINF的实验研究

国家自然科学基金

0+阅读 · 2008年12月31日

Fairness-guided Few-shot Prompting for Large Language Models

Arxiv

0+阅读 · 2023年3月23日

Towards Better Dynamic Graph Learning: New Architecture and Unified Library

Arxiv

0+阅读 · 2023年3月23日

Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Arxiv

0+阅读 · 2023年3月22日

eP-ALM: Efficient Perceptual Augmentation of Language Models

Arxiv

0+阅读 · 2023年3月20日

Pre-Trained Models: Past, Present and Future

Arxiv

19+阅读 · 2021年6月15日

A continual learning survey: Defying forgetting in classification tasks

Arxiv

32+阅读 · 2021年4月16日

Towards Open World Object Detection

Arxiv

13+阅读 · 2021年3月3日

Making Pre-trained Language Models Better Few-shot Learners

Arxiv

14+阅读 · 2020年12月31日

K-BERT: Enabling Language Representation with Knowledge Graph

K-BERT: Enabling Language Representation with Knowledge Graph

Arxiv

19+阅读 · 2019年9月17日

Weakly Supervised One-Shot Detection with Attention Siamese Networks

Arxiv

14+阅读 · 2018年1月12日

VIP会员

文章信息

相关主题

语言模型化

小样本学习

词元分析器

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

机器学习组合优化

机器学习组合优化

专知会员服务

110+阅读 · 2021年2月16日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

操作系统智能体：基于多模态大模型（MLLM）的通用计算设备智能体综述

《美国太空军系统全生命周期建模、仿真与分析效能提升方案》最新84页报告

【博士论文】推进数据高效的深度学习：非参数 Transformer、主动测试与上下文学习

自主人工智能：未来战争是否将是自主化的？

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

征稿 | CFP：Special Issue of NLP and KG(JCR Q2，IF2.67)

开放知识图谱

1+阅读 · 2022年4月4日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

BERT/注意力机制/Transformer/迁移学习NLP资源大列表：awesome-bert-nlp

AINLP

40+阅读 · 2019年6月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

相关论文

Fairness-guided Few-shot Prompting for Large Language Models

Arxiv

0+阅读 · 2023年3月23日

Towards Better Dynamic Graph Learning: New Architecture and Unified Library

Arxiv

0+阅读 · 2023年3月23日

Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Arxiv

0+阅读 · 2023年3月22日

eP-ALM: Efficient Perceptual Augmentation of Language Models

Arxiv

0+阅读 · 2023年3月20日

Pre-Trained Models: Past, Present and Future

Arxiv

19+阅读 · 2021年6月15日

A continual learning survey: Defying forgetting in classification tasks

Arxiv

32+阅读 · 2021年4月16日

Towards Open World Object Detection

Arxiv

13+阅读 · 2021年3月3日

Making Pre-trained Language Models Better Few-shot Learners

Arxiv

14+阅读 · 2020年12月31日

K-BERT: Enabling Language Representation with Knowledge Graph

K-BERT: Enabling Language Representation with Knowledge Graph

Arxiv

19+阅读 · 2019年9月17日

Weakly Supervised One-Shot Detection with Attention Siamese Networks

Arxiv

14+阅读 · 2018年1月12日

相关基金

胡杨耐盐的表观基因组学研究

国家自然科学基金

0+阅读 · 2015年12月31日

茉莉酸调控水稻花器官发育分子调控网络的研究

国家自然科学基金

0+阅读 · 2014年12月31日

自噬对高脂膳食诱导的血管内皮细胞损伤的保护作用及分子机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

LPA-LPAR1信号通路调控放射性肺损伤的机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

Rbp1在特络细胞联合骨髓间充质干细胞治疗急性肺损伤中的调节作用和分子机制

国家自然科学基金

0+阅读 · 2014年12月31日

BRCA1蛋白出核的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

混凝土细观损伤模拟与数值尺寸效应研究

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预NF-кB信号通路防治动脉粥样硬化的作用及机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

遍历哈密顿系统的谱理论

国家自然科学基金

0+阅读 · 2009年12月31日

C3G－SOD双靶点阻遏ROS多环节改善PINF的实验研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员