与违规学习一起不受监督的浓情信息检索 (Unsupervised Dense Information Retrieval with Contrastive Learning) - 专知论文

会员服务 ·

0

无监督 · contrastive · INFORMS · BM25 · 对比学习 ·

2022 年 5 月 26 日

Unsupervised Dense Information Retrieval with Contrastive Learning

翻译：与违规学习一起不受监督的浓情信息检索

Gautier Izacard,Mathilde Caron,Lucas Hosseini,Sebastian Riedel,Piotr Bojanowski,Armand Joulin,Edouard Grave

Recently, information retrieval has seen the emergence of dense retrievers, based on neural networks, as an alternative to classical sparse methods based on term-frequency. These models have obtained state-of-the-art results on datasets and tasks where large training sets are available. However, they do not transfer well to new applications with no training data, and are outperformed by unsupervised term-frequency methods such as BM25. In this work, we explore the limits of contrastive learning as a way to train unsupervised dense retrievers and show that it leads to strong performance in various retrieval settings. On the BEIR benchmark our unsupervised model outperforms BM25 on 11 out of 15 datasets for the Recall@100 metric. When used as pre-training before fine-tuning, either on a few thousands in-domain examples or on the large MS MARCO dataset, our contrastive model leads to improvements on the BEIR benchmark. Finally, we evaluate our approach for multi-lingual retrieval, where training data is even scarcer than for English, and show that our approach leads to strong unsupervised performance. Our model also exhibits strong cross-lingual transfer when fine-tuned on supervised English data only and evaluated on low resources language such as Swahili. We show that our unsupervised models can perform cross-lingual retrieval between different scripts, such as retrieving English documents from Arabic queries, which would not be possible with term matching methods.

翻译：最近,信息检索发现,在神经网络的基础上出现了密密的检索器,以神经网络为基础,作为基于使用频率的经典稀疏方法的替代方法。这些模型在有大型培训数据集的情况下,在数据集和任务方面获得了最先进的结果。然而,这些模型没有很好地向没有培训数据的新应用程序转移,而且以未受监督的术语频率方法,如BM25等,其表现优于未受监督的术语检索器。在这项工作中,我们探索对比学习的局限性,以此作为培训不受监督的密集检索器的一种方法,并表明它导致各种检索环境中的强效业绩。在BEIR基准中,这些模型测量了我们未经监督的模型在15个数据集中,11个模型是BMB25。在进行微调前的训练时,没有很好地将BMBMARCO数据集(例如BM25),我们的对比模型导致BER基准的改进。最后,我们评估我们多语言检索方法的方法,即培训数据比英语更稀少,并且显示我们的方法是强的、强的BBM25,在检索过程中,我们的方法是强的不精准的、不精细的英文本检索。

0

相关内容

无监督

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

19篇ICML2019论文摘录选读！

19篇ICML2019论文摘录选读！

专知

28+阅读 · 2019年4月28日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

磁性纳米团簇表面贵金属纳米粒子的分散稳定机制和催化性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

MOFs纳米粒子的制备及其对不相容共混物相结构的调控与稳定作用

国家自然科学基金

0+阅读 · 2015年12月31日

精确控制液膜破裂与纳米粒子图案化组装研究

国家自然科学基金

0+阅读 · 2014年12月31日

考虑非定常气动力随机不确定性的气动弹性研究

国家自然科学基金

0+阅读 · 2013年12月31日

空间高稳定精密跟瞄Stewart平台设计优化与控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

P2P-Grid环境中分布式不确定本体模型的研究

国家自然科学基金

0+阅读 · 2013年12月31日

大数据环境下面向科学研究第四范式的信息资源云研究

国家自然科学基金

1+阅读 · 2012年12月31日

多铁性LSCMO/PMN-PT磁电复合薄膜的制备、表征及原型器件探索

国家自然科学基金

0+阅读 · 2012年12月31日

超临界CO2脉冲电沉积钴基纳米合金薄膜的形成机理、结构控制及摩擦磨损行为研究

国家自然科学基金

0+阅读 · 2012年12月31日

SrxBa1-xNb2O6纳米陶瓷与薄膜的电卡效应

国家自然科学基金

0+阅读 · 2012年12月31日

Continual Contrastive Learning for Image Classification

Arxiv

0+阅读 · 2022年7月14日

DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval

Arxiv

0+阅读 · 2022年7月13日

Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning

Arxiv

0+阅读 · 2022年7月12日

Pre-training Methods in Information Retrieval

Arxiv

16+阅读 · 2021年11月27日

Sequence Level Contrastive Learning for Text Summarization

Sequence Level Contrastive Learning for Text Summarization

Arxiv

14+阅读 · 2021年9月24日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

A Decade Survey of Content Based Image Retrieval using Deep Learning

Arxiv

23+阅读 · 2020年11月23日

PROP: Pre-training with Representative Words Prediction for Ad-hoc Retrieval

Arxiv

11+阅读 · 2020年10月20日

A survey on deep hashing for image retrieval

A survey on deep hashing for image retrieval

Arxiv

15+阅读 · 2020年6月10日

Graph Convolutional Networks for Text Classification

Arxiv

11+阅读 · 2018年10月17日

VIP会员

文章信息

相关主题

相关VIP内容

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

80+阅读 · 2020年7月26日

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

【干货书】真实机器学习，264页pdf，Real-World Machine Learning

专知会员服务

115+阅读 · 2020年4月5日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

95+阅读 · 2020年3月12日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

181+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

《生成式人工智能与大/小语言模型在供应链管理决策优化与可持续性提升中的作用评估》最新51页

白宫发布《赢得AI竞赛：美国人工智能行动计划》最新28页

地下战：地下空间的战略博弈

《美地下作战条令手册》228页

相关资讯

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Plenary Talk1

【ICIG2021】Latest News & Announcements of the Plenary Talk1

中国图象图形学学会CSIG

0+阅读 · 2021年11月1日

【ICIG2021】Latest News & Announcements of the Industry Talk2

【ICIG2021】Latest News & Announcements of the Industry Talk2

中国图象图形学学会CSIG

0+阅读 · 2021年7月29日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知

133+阅读 · 2020年3月18日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

19篇ICML2019论文摘录选读！

19篇ICML2019论文摘录选读！

专知

28+阅读 · 2019年4月28日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

vae 相关论文表示学习 1

vae 相关论文表示学习 1

CreateAMind

12+阅读 · 2018年9月6日

相关论文

Continual Contrastive Learning for Image Classification

Arxiv

0+阅读 · 2022年7月14日

DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval

Arxiv

0+阅读 · 2022年7月13日

Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning

Arxiv

0+阅读 · 2022年7月12日

Pre-training Methods in Information Retrieval

Arxiv

16+阅读 · 2021年11月27日

Sequence Level Contrastive Learning for Text Summarization

Sequence Level Contrastive Learning for Text Summarization

Arxiv

14+阅读 · 2021年9月24日

Deep Image Retrieval: A Survey

Arxiv

16+阅读 · 2021年1月27日

A Decade Survey of Content Based Image Retrieval using Deep Learning

Arxiv

23+阅读 · 2020年11月23日

PROP: Pre-training with Representative Words Prediction for Ad-hoc Retrieval

Arxiv

11+阅读 · 2020年10月20日

A survey on deep hashing for image retrieval

A survey on deep hashing for image retrieval

Arxiv

15+阅读 · 2020年6月10日

Graph Convolutional Networks for Text Classification

Arxiv

11+阅读 · 2018年10月17日

相关基金

磁性纳米团簇表面贵金属纳米粒子的分散稳定机制和催化性能研究

国家自然科学基金

0+阅读 · 2015年12月31日

MOFs纳米粒子的制备及其对不相容共混物相结构的调控与稳定作用

国家自然科学基金

0+阅读 · 2015年12月31日

精确控制液膜破裂与纳米粒子图案化组装研究

国家自然科学基金

0+阅读 · 2014年12月31日

考虑非定常气动力随机不确定性的气动弹性研究

国家自然科学基金

0+阅读 · 2013年12月31日

空间高稳定精密跟瞄Stewart平台设计优化与控制方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

P2P-Grid环境中分布式不确定本体模型的研究

国家自然科学基金

0+阅读 · 2013年12月31日

大数据环境下面向科学研究第四范式的信息资源云研究

国家自然科学基金

1+阅读 · 2012年12月31日

多铁性LSCMO/PMN-PT磁电复合薄膜的制备、表征及原型器件探索

国家自然科学基金

0+阅读 · 2012年12月31日

超临界CO2脉冲电沉积钴基纳米合金薄膜的形成机理、结构控制及摩擦磨损行为研究

国家自然科学基金

0+阅读 · 2012年12月31日

SrxBa1-xNb2O6纳米陶瓷与薄膜的电卡效应

国家自然科学基金

0+阅读 · 2012年12月31日

微信扫码咨询专知VIP会员