将内文学习解释为隐含的贝耶斯推论 (An Explanation of In-context Learning as Implicit Bayesian Inference) - 专知论文

会员服务 ·

0

Learning · Prompt · 贝叶斯推断 · 推断 · 样例 ·

2022 年 7 月 21 日

An Explanation of In-context Learning as Implicit Bayesian Inference

翻译：将内文学习解释为隐含的贝耶斯推论

Sang Michael Xie,Aditi Raghunathan,Percy Liang,Tengyu Ma

from arxiv, ICLR 2022

Large language models (LMs) such as GPT-3 have the surprising ability to do in-context learning, where the model learns to do a downstream task simply by conditioning on a prompt consisting of input-output examples. The LM learns from these examples without being explicitly pretrained to learn. Thus, it is unclear what enables in-context learning. In this paper, we study how in-context learning can emerge when pretraining documents have long-range coherence. Here, the LM must infer a latent document-level concept to generate coherent next tokens during pretraining. At test time, in-context learning occurs when the LM also infers a shared latent concept between examples in a prompt. We prove when this occurs despite a distribution mismatch between prompts and pretraining data in a setting where the pretraining distribution is a mixture of HMMs. In contrast to messy large-scale datasets used to train LMs capable of in-context learning, we generate a small-scale synthetic dataset (GINC) where Transformers and LSTMs both exhibit in-context learning. Beyond the theory, experiments on GINC exhibit large-scale real-world phenomena including improved in-context performance with model scaling (despite the same pretraining loss), sensitivity to example order, and instances where zero-shot is better than few-shot in-context learning.

翻译：GPT-3等大型语言模型(LMS)具有令人惊讶的在文字上学习的能力,而该模型仅靠由投入产出实例组成的快速范例来学习,就学会了下游任务。LMS从这些实例中学习,而没有经过明确的培训学习。因此,不清楚是什么使得在文字上学习。在本文中,我们研究在训练前文件具有长期一致性时,如何出现在文字上学习。在这里,LM必须推导一种潜在的文件级概念,以便在培训前产生一致的下一个标志。在测试时,当LM还推介一个快速实例之间的共同潜在概念时,就会发生文字上学习。我们证明,尽管在培训前的分布是HMMM的零混合的环境下,在这种环境中,在提示和预培训前的数据之间分配不匹配。与培训前用于培训LMS的大规模数据集相比,我们产生了一种小规模的合成数据集(GINC),其中变换器和LSTMS都出现在文字上很少的模型中,除了理论上、在GINC系统上进行更好的学习外,在理论上进行更好的实验,在实际损失顺序上进行更好的实验。

0

相关内容

Learning

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【医学图像处理中的因果性】52页ppt，Causality Matters in Medical Imaging

【医学图像处理中的因果性】52页ppt，Causality Matters in Medical Imaging

专知会员服务

60+阅读 · 2020年3月14日

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

专知会员服务

82+阅读 · 2020年2月27日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

Multi-Task Learning的几篇综述文章

Multi-Task Learning的几篇综述文章

深度学习自然语言处理

15+阅读 · 2020年6月15日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

基于IIM模型的城市关联基础设施系统的脆弱性与弹性评价研究

国家自然科学基金

1+阅读 · 2015年12月31日

外泌体（Exosome）在小肠上皮损伤修复的作用机制及甘草的干预研究

国家自然科学基金

0+阅读 · 2014年12月31日

可压缩湍流粒子输运的拉格朗日（Lagrangian）研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于大涡模拟的风致超高层建筑复杂运动数值模拟方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

含极性非质子溶剂的离子液体-kosmotropic盐双水相体系的研究

国家自然科学基金

0+阅读 · 2012年12月31日

TREM-1/DAP12/ NF-κB信号通路在6-姜烯酚抗动脉粥样硬化中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

LIMK1：罗格列酮抑制人胃癌细胞增殖、迁移及侵袭的作用靶点

国家自然科学基金

0+阅读 · 2012年12月31日

石榴鞣花酸调控胆固醇代谢的分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

BEC的保几何结构数值模拟与研究

国家自然科学基金

0+阅读 · 2011年12月31日

西南岩溶裂隙-管道介质地下水流运动规律试验研究

国家自然科学基金

0+阅读 · 2011年12月31日

Learning to Answer Semantic Queries over Code

Arxiv

0+阅读 · 2022年9月17日

Conformal prediction beyond exchangeability

Arxiv

0+阅读 · 2022年9月16日

Dataset Inference for Self-Supervised Models

Arxiv

0+阅读 · 2022年9月16日

Maximum Likelihood Training of Implicit Nonlinear Diffusion Models

Arxiv

0+阅读 · 2022年9月16日

On the Relation between Sensitivity and Accuracy in In-context Learning

Arxiv

0+阅读 · 2022年9月16日

Robust explicit estimation of the log-logistic distribution with applications

Arxiv

0+阅读 · 2022年9月15日

On the detrimental effect of invariances in the likelihood for variational inference

Arxiv

0+阅读 · 2022年9月15日

From Easy to Hard: A Dual Curriculum Learning Framework for Context-Aware Document Ranking

Arxiv

0+阅读 · 2022年9月15日

Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

Arxiv

10+阅读 · 2022年2月10日

Generative Models as a Data Source for Multiview Representation Learning

Arxiv

16+阅读 · 2021年6月9日

VIP会员

文章信息

相关主题

贝叶斯推断

相关VIP内容

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【医学图像处理中的因果性】52页ppt，Causality Matters in Medical Imaging

【医学图像处理中的因果性】52页ppt，Causality Matters in Medical Imaging

专知会员服务

60+阅读 · 2020年3月14日

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

【经典书】数据挖掘：理论、算法与示例，347页pdf，Nong Ye，Arizona State University

专知会员服务

82+阅读 · 2020年2月27日

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

【新书：机器学习简介】《A Concise Introduction to Machine Learning》by A.C. Faul (CRC 2019)

专知会员服务

77+阅读 · 2020年2月8日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

操作系统智能体：基于多模态大模型（MLLM）的通用计算设备智能体综述

《美国太空军系统全生命周期建模、仿真与分析效能提升方案》最新84页报告

【博士论文】推进数据高效的深度学习：非参数 Transformer、主动测试与上下文学习

自主人工智能：未来战争是否将是自主化的？

相关资讯

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

Multi-Task Learning的几篇综述文章

Multi-Task Learning的几篇综述文章

深度学习自然语言处理

15+阅读 · 2020年6月15日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

Learning to Answer Semantic Queries over Code

Arxiv

0+阅读 · 2022年9月17日

Conformal prediction beyond exchangeability

Arxiv

0+阅读 · 2022年9月16日

Dataset Inference for Self-Supervised Models

Arxiv

0+阅读 · 2022年9月16日

Maximum Likelihood Training of Implicit Nonlinear Diffusion Models

Arxiv

0+阅读 · 2022年9月16日

On the Relation between Sensitivity and Accuracy in In-context Learning

Arxiv

0+阅读 · 2022年9月16日

Robust explicit estimation of the log-logistic distribution with applications

Arxiv

0+阅读 · 2022年9月15日

On the detrimental effect of invariances in the likelihood for variational inference

Arxiv

0+阅读 · 2022年9月15日

From Easy to Hard: A Dual Curriculum Learning Framework for Context-Aware Document Ranking

Arxiv

0+阅读 · 2022年9月15日

Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks

Arxiv

10+阅读 · 2022年2月10日

Generative Models as a Data Source for Multiview Representation Learning

Arxiv

16+阅读 · 2021年6月9日

相关基金

基于IIM模型的城市关联基础设施系统的脆弱性与弹性评价研究

国家自然科学基金

1+阅读 · 2015年12月31日

外泌体（Exosome）在小肠上皮损伤修复的作用机制及甘草的干预研究

国家自然科学基金

0+阅读 · 2014年12月31日

可压缩湍流粒子输运的拉格朗日（Lagrangian）研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于大涡模拟的风致超高层建筑复杂运动数值模拟方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

含极性非质子溶剂的离子液体-kosmotropic盐双水相体系的研究

国家自然科学基金

0+阅读 · 2012年12月31日

TREM-1/DAP12/ NF-κB信号通路在6-姜烯酚抗动脉粥样硬化中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

LIMK1：罗格列酮抑制人胃癌细胞增殖、迁移及侵袭的作用靶点

国家自然科学基金

0+阅读 · 2012年12月31日

石榴鞣花酸调控胆固醇代谢的分子机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

BEC的保几何结构数值模拟与研究

国家自然科学基金

0+阅读 · 2011年12月31日

西南岩溶裂隙-管道介质地下水流运动规律试验研究

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员