AugTriiever: 由可缩放数据放大器进行不受监督的常量检索</s> (AugTriever: Unsupervised Dense Retrieval by Scalable Data Augmentation) - 专知论文

会员服务 ·

0

无监督 · Performer · MoDELS · 数据增强 · Extensibility ·

2023 年 3 月 7 日

AugTriever: Unsupervised Dense Retrieval by Scalable Data Augmentation

翻译：AugTriiever: 由可缩放数据放大器进行不受监督的常量检索

Rui Meng,Ye Liu,Semih Yavuz,Divyansh Agarwal,Lifu Tu,Ning Yu,Jianguo Zhang,Meghana Bhat,Yingbo Zhou

Dense retrievers have made significant strides in text retrieval and open-domain question answering, even though most achievements were made possible only with large amounts of human supervision. In this work, we aim to develop unsupervised methods by proposing two methods that create pseudo query-document pairs and train dense retrieval models in an annotation-free and scalable manner: query extraction and transferred query generation. The former method produces pseudo queries by selecting salient spans from the original document. The latter utilizes generation models trained for other NLP tasks (e.g., summarization) to produce pseudo queries. Extensive experiments show that models trained with the proposed augmentation methods can perform comparably well (or better) to multiple strong baselines. Combining those strategies leads to further improvements, achieving the state-of-the-art performance of unsupervised dense retrieval on both BEIR and ODQA datasets.

翻译：大量检索者在文本检索和开放域问题解答方面取得了长足的进步,尽管大多数成就只有在大量的人力监督下才有可能实现。在这项工作中,我们的目标是制定不受监督的方法,提出两种方法来创建假的查询文档配对,并以无注释和可缩放的方式培训密集检索模型:查询提取和传输查询生成。前一种方法通过从原始文档中选择突出的空格产生伪查询。后一种方法利用经过培训的用于其他国家实验室任务(如汇总)的生成模型来生成假查询。广泛的实验表明,经过培训的增强方法模型能够很好(或更好)地运行到多个强大的基线。合并这些战略可以带来进一步的改进,在BEIR和ODQA数据集上实现无监控密度检索的最先进性能。</s>

0

相关内容

无监督

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

最新5篇生成对抗网络相关论文推荐—FusedGAN、DeblurGAN、AdvGAN、CipherGAN、MMD GANS

最新5篇生成对抗网络相关论文推荐—FusedGAN、DeblurGAN、AdvGAN、CipherGAN、MMD GANS

专知

23+阅读 · 2018年1月18日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

【推荐】深度学习目标检测全面综述

【推荐】深度学习目标检测全面综述

机器学习研究会

21+阅读 · 2017年9月13日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

基于自噬系统mTOR信号通路探讨扶正祛邪中药小复方干预阿尔茨海默病模型的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

CIP2A对蛋白磷酸酯酶2A的调节及其在阿尔茨海默病发病中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

Runx3基因DNA甲基化介导BPD肺上皮细胞转分化的作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

丹参经表观遗传调控Nrf2/ARE通路及降低核苷酸类似物肾毒性的作用机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于EEG和fNIRS的多模态脑机接口运动想象参数研究

国家自然科学基金

1+阅读 · 2012年12月31日

基于CYP450酶表达调控及代谢组学的五味子醋制保肝作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

P53蛋白调节mTOR信号通路诱导胰腺癌吉西他滨耐药的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

炎症细胞因子DNA甲基化影响炎性衰老的机制

国家自然科学基金

0+阅读 · 2011年12月31日

DNA甲基化介导的CLDN6表达沉默机制及其对人乳腺癌细胞转移表型的影响

国家自然科学基金

0+阅读 · 2011年12月31日

TGF-βsmads信号通路对失神经骨骼肌纤维化调控机制的实验研究

国家自然科学基金

0+阅读 · 2008年12月31日

A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt Learning

Arxiv

0+阅读 · 2023年4月28日

Multivariate Representation Learning for Information Retrieval

Arxiv

0+阅读 · 2023年4月27日

Person Re-ID through Unsupervised Hypergraph Rank Selection and Fusion

Arxiv

0+阅读 · 2023年4月27日

Large Language Models are Strong Zero-Shot Retriever

Arxiv

0+阅读 · 2023年4月27日

Retrieval-based Knowledge Augmented Vision Language Pre-training

Arxiv

0+阅读 · 2023年4月27日

A Personalized Dense Retrieval Framework for Unified Information Access

A Personalized Dense Retrieval Framework for Unified Information Access

Arxiv

0+阅读 · 2023年4月26日

ContrastMask: Contrastive Learning to Segment Every Thing

Arxiv

15+阅读 · 2022年3月18日

MetAug: Contrastive Learning via Meta Feature Augmentation

Arxiv

10+阅读 · 2022年3月10日

Unifying Vision-and-Language Tasks via Text Generation

Arxiv

10+阅读 · 2021年2月4日

On Feature Normalization and Data Augmentation

On Feature Normalization and Data Augmentation

Arxiv

15+阅读 · 2020年2月25日

VIP会员

文章信息

相关主题

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【Google】深度学习对抗鲁棒性，43页ppt

专知会员服务

45+阅读 · 2020年10月31日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津博士论文】零样本强化学习综述

《美军条令：陆军指挥官与规划人员地理空间指南》60页

战术边缘指挥控制：防务面临的核心挑战

迈向开放世界检测：综述

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

最新5篇生成对抗网络相关论文推荐—FusedGAN、DeblurGAN、AdvGAN、CipherGAN、MMD GANS

最新5篇生成对抗网络相关论文推荐—FusedGAN、DeblurGAN、AdvGAN、CipherGAN、MMD GANS

专知

23+阅读 · 2018年1月18日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】YOLO实时目标检测(6fps)

【推荐】YOLO实时目标检测(6fps)

机器学习研究会

20+阅读 · 2017年11月5日

【推荐】深度学习目标检测全面综述

【推荐】深度学习目标检测全面综述

机器学习研究会

21+阅读 · 2017年9月13日

【推荐】GAN架构入门综述(资源汇总)

【推荐】GAN架构入门综述(资源汇总)

机器学习研究会

10+阅读 · 2017年9月3日

相关论文

A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt Learning

Arxiv

0+阅读 · 2023年4月28日

Multivariate Representation Learning for Information Retrieval

Arxiv

0+阅读 · 2023年4月27日

Person Re-ID through Unsupervised Hypergraph Rank Selection and Fusion

Arxiv

0+阅读 · 2023年4月27日

Large Language Models are Strong Zero-Shot Retriever

Arxiv

0+阅读 · 2023年4月27日

Retrieval-based Knowledge Augmented Vision Language Pre-training

Arxiv

0+阅读 · 2023年4月27日

A Personalized Dense Retrieval Framework for Unified Information Access

A Personalized Dense Retrieval Framework for Unified Information Access

Arxiv

0+阅读 · 2023年4月26日

ContrastMask: Contrastive Learning to Segment Every Thing

Arxiv

15+阅读 · 2022年3月18日

MetAug: Contrastive Learning via Meta Feature Augmentation

Arxiv

10+阅读 · 2022年3月10日

Unifying Vision-and-Language Tasks via Text Generation

Arxiv

10+阅读 · 2021年2月4日

On Feature Normalization and Data Augmentation

On Feature Normalization and Data Augmentation

Arxiv

15+阅读 · 2020年2月25日

相关基金

基于自噬系统mTOR信号通路探讨扶正祛邪中药小复方干预阿尔茨海默病模型的机制研究

国家自然科学基金

0+阅读 · 2015年12月31日

CIP2A对蛋白磷酸酯酶2A的调节及其在阿尔茨海默病发病中的作用

国家自然科学基金

0+阅读 · 2014年12月31日

Runx3基因DNA甲基化介导BPD肺上皮细胞转分化的作用及机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

丹参经表观遗传调控Nrf2/ARE通路及降低核苷酸类似物肾毒性的作用机制研究

国家自然科学基金

0+阅读 · 2014年12月31日

基于EEG和fNIRS的多模态脑机接口运动想象参数研究

国家自然科学基金

1+阅读 · 2012年12月31日

基于CYP450酶表达调控及代谢组学的五味子醋制保肝作用机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

P53蛋白调节mTOR信号通路诱导胰腺癌吉西他滨耐药的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

炎症细胞因子DNA甲基化影响炎性衰老的机制

国家自然科学基金

0+阅读 · 2011年12月31日

DNA甲基化介导的CLDN6表达沉默机制及其对人乳腺癌细胞转移表型的影响

国家自然科学基金

0+阅读 · 2011年12月31日

TGF-βsmads信号通路对失神经骨骼肌纤维化调控机制的实验研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员