他们知道什么? 他们知道什么? 他们怎么知道? (What do tokens know about their characters and how do they know it?) - 专知论文

会员服务 ·

0

INFORMS · 词元分析器 · MoDELS · Alphabet · 语言模型化 ·

2022 年 6 月 6 日

What do tokens know about their characters and how do they know it?

翻译：他们知道什么? 他们知道什么? 他们怎么知道?

Ayush Kaushal,Kyle Mahowald

Pre-trained language models (PLMs) that use subword tokenization schemes can succeed at a variety of language tasks that require character-level information, despite lacking explicit access to the character composition of tokens. Here, studying a range of models (e.g., GPT- J, BERT, RoBERTa, GloVe), we probe what word pieces encode about character-level information by training classifiers to predict the presence or absence of a particular alphabetical character in a token, based on its embedding (e.g., probing whether the model embedding for "cat" encodes that it contains the character "a"). We find that these models robustly encode character-level information and, in general, larger models perform better at the task. We show that these results generalize to characters from non-Latin alphabets (Arabic, Devanagari, and Cyrillic). Then, through a series of experiments and analyses, we investigate the mechanisms through which PLMs acquire English-language character information during training and argue that this knowledge is acquired through multiple phenomena, including a systematic relationship between particular characters and particular parts of speech, as well as natural variability in the tokenization of related strings.

翻译：使用子词符号化办法的经过事先训练的语言模型(PLM)在各种语言任务中可以取得成功,这些语言任务需要品格级信息,尽管缺乏明确获取象征品的品格组成。在这里,我们研究一系列模型(例如,GPT-J、BERT、ROBERTA、GloVe),我们通过培训分类人员,根据嵌入方式(例如,检验是否嵌入包含“a”特性的“Cat”编码的模型),可以成功完成各种需要品格级信息的语文任务。我们发现,这些模型强有力地编码了字符级信息,总体而言,较大的模型在执行任务时表现更好。我们表明,这些结果概括了非拉丁字母(阿拉伯文、德瓦纳加里和西里尔利奇)的字符。然后,通过一系列试验和分析,我们调查PLMs在培训期间获取英语特征信息的机制,并论证这种知识是通过多种现象获得的,包括特定字符和特定语言部分之间的系统关系,以及象征质化的自然变异性。

0

相关内容

INFORMS

《计算机信息》杂志发表高质量的论文，扩大了运筹学和计算的范围，寻求有关理论、方法、实验、系统和应用方面的原创研究论文、新颖的调查和教程论文，以及描述新的和有用的软件工具的论文。官网链接：https://pubsonline.informs.org/journal/ijoc

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

53+阅读 · 2021年1月20日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

有限域上多项式的p-进与T-进指数和

国家自然科学基金

0+阅读 · 2013年12月31日

BER通路基因miRNA结合位点基因多态性与结直肠癌易感性的关联及功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

miRNA-92a对Rho激酶调控的动脉粥样硬化血管重构的影响及机制

国家自然科学基金

0+阅读 · 2013年12月31日

Doublesex基因在对虾性别决定和分化中的功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

过渡金属簇共轭桥联稀土近红外发光功能配合物的组装、结构与性能

国家自然科学基金

0+阅读 · 2012年12月31日

同型半胱氨酸致动脉粥样硬化中"c-myc/miRNAs/FABP4"交互作用分子网络的构建及潜在干预靶位的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于pQDs-CD133mAb的双模态探针对胶质瘤CD133+干细胞靶向成像的实验研究

国家自然科学基金

0+阅读 · 2011年12月31日

LRRC4-AP2/SP1-miR182-LRRC4和LRRC4-miR-185-DNMT1-LRRC4调控环路在脑胶质瘤中相互调控的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

胶质瘤中AKT/βatenin信号转导通路转录调控miR-21的新机制

国家自然科学基金

0+阅读 · 2009年12月31日

Transforming Wikipedia into Augmented Data for Query-Focused Summarization

Arxiv

0+阅读 · 2022年7月22日

Widespread Underestimation of Sensitivity in Differentially Private Libraries and How to Fix It

Widespread Underestimation of Sensitivity in Differentially Private Libraries and How to Fix It

Arxiv

0+阅读 · 2022年7月21日

Incentive Designs for Stackelberg Games with a Large Number of Followers and their Mean-Field Limits

Incentive Designs for Stackelberg Games with a Large Number of Followers and their Mean-Field Limits

Arxiv

0+阅读 · 2022年7月21日

Algorithmic encoding of protected characteristics in image-based models for disease detection

Algorithmic encoding of protected characteristics in image-based models for disease detection

Arxiv

0+阅读 · 2022年7月21日

Image and Model Transformation with Secret Key for Vision Transformer

Arxiv

0+阅读 · 2022年7月21日

End-to-End and Self-Supervised Learning for ComParE 2022 Stuttering Sub-Challenge

Arxiv

0+阅读 · 2022年7月20日

Pre-Trained Models: Past, Present and Future

Arxiv

19+阅读 · 2021年6月15日

HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions

HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions

Arxiv

10+阅读 · 2020年12月31日

A Primer in BERTology: What we know about how BERT works

A Primer in BERTology: What we know about how BERT works

Arxiv

34+阅读 · 2020年2月27日

Representation Learning with Ordered Relation Paths for Knowledge Graph Completion

Representation Learning with Ordered Relation Paths for Knowledge Graph Completion

Arxiv

12+阅读 · 2019年9月26日

VIP会员

文章信息

相关主题

词元分析器

语言模型化

相关VIP内容

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

剑桥大学《数据科学: 原理与实践》课程，附PPT下载

专知会员服务

53+阅读 · 2021年1月20日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

2020数据工程师成长路线图

专知会员服务

41+阅读 · 2020年9月6日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

人工智能治理的未来

模态感知的特征匹配：单一模态与跨模态技术的全面综述

无监督行人重识别研究综述

【牛津博士论文】面向神经影像应用的可扩展且可解释的空间模型

相关资讯

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Tutorial

【ICIG2021】Latest News & Announcements of the Tutorial

中国图象图形学学会CSIG

3+阅读 · 2021年12月20日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium3

中国图象图形学学会CSIG

0+阅读 · 2021年11月9日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium2

中国图象图形学学会CSIG

0+阅读 · 2021年11月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Transforming Wikipedia into Augmented Data for Query-Focused Summarization

Arxiv

0+阅读 · 2022年7月22日

Widespread Underestimation of Sensitivity in Differentially Private Libraries and How to Fix It

Widespread Underestimation of Sensitivity in Differentially Private Libraries and How to Fix It

Arxiv

0+阅读 · 2022年7月21日

Incentive Designs for Stackelberg Games with a Large Number of Followers and their Mean-Field Limits

Incentive Designs for Stackelberg Games with a Large Number of Followers and their Mean-Field Limits

Arxiv

0+阅读 · 2022年7月21日

Algorithmic encoding of protected characteristics in image-based models for disease detection

Algorithmic encoding of protected characteristics in image-based models for disease detection

Arxiv

0+阅读 · 2022年7月21日

Image and Model Transformation with Secret Key for Vision Transformer

Arxiv

0+阅读 · 2022年7月21日

End-to-End and Self-Supervised Learning for ComParE 2022 Stuttering Sub-Challenge

Arxiv

0+阅读 · 2022年7月20日

Pre-Trained Models: Past, Present and Future

Arxiv

19+阅读 · 2021年6月15日

HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions

HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions

Arxiv

10+阅读 · 2020年12月31日

A Primer in BERTology: What we know about how BERT works

A Primer in BERTology: What we know about how BERT works

Arxiv

34+阅读 · 2020年2月27日

Representation Learning with Ordered Relation Paths for Knowledge Graph Completion

Representation Learning with Ordered Relation Paths for Knowledge Graph Completion

Arxiv

12+阅读 · 2019年9月26日

相关基金

miRNAs调控柿单宁合成代谢机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

有限域上多项式的p-进与T-进指数和

国家自然科学基金

0+阅读 · 2013年12月31日

BER通路基因miRNA结合位点基因多态性与结直肠癌易感性的关联及功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

miRNA-92a对Rho激酶调控的动脉粥样硬化血管重构的影响及机制

国家自然科学基金

0+阅读 · 2013年12月31日

Doublesex基因在对虾性别决定和分化中的功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

过渡金属簇共轭桥联稀土近红外发光功能配合物的组装、结构与性能

国家自然科学基金

0+阅读 · 2012年12月31日

同型半胱氨酸致动脉粥样硬化中"c-myc/miRNAs/FABP4"交互作用分子网络的构建及潜在干预靶位的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于pQDs-CD133mAb的双模态探针对胶质瘤CD133+干细胞靶向成像的实验研究

国家自然科学基金

0+阅读 · 2011年12月31日

LRRC4-AP2/SP1-miR182-LRRC4和LRRC4-miR-185-DNMT1-LRRC4调控环路在脑胶质瘤中相互调控的机制研究

国家自然科学基金

0+阅读 · 2011年12月31日

胶质瘤中AKT/βatenin信号转导通路转录调控miR-21的新机制

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员