The translated title (The Archive Query Log: Mining Millions of Search Result Pages of Hundreds of Search Engines from 25 Years of Web Archives) - 专知论文

会员服务 ·

0

查询日志 · 搜索 · 搜索引擎 · 引擎 · 检索模型 ·

2023 年 4 月 2 日

The Archive Query Log: Mining Millions of Search Result Pages of Hundreds of Search Engines from 25 Years of Web Archives

翻译：The translated title

Jan Heinrich Reimer,Sebastian Schmidt,Maik Fröbe,Lukas Gienapp,Harrisen Scells,Benno Stein,Matthias Hagen,Martin Potthast

from arxiv, 12 pages. To be published in the proceedings of SIGIR 2023

The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 166 million search result pages, and 1.7 billion search results across 550 search providers. Although many query logs have been studied in the literature, the search providers that own them generally do not publish their logs to protect user privacy and vital business data. Of the few query logs publicly available, none combines size, scope, and diversity. The AQL is the first to do so, enabling research on new retrieval models and (diachronic) search engine analyses. Provided in a privacy-preserving manner, it promotes open research as well as more transparency and accountability in the search industry.

翻译：网络档案查询日志：挖掘25年来550个搜索引擎数百万个搜索结果页面的数据 The translated abstract 网络档案查询日志（AQL）是在互联网档案馆收集的以前未使用的全面查询日志，历时25年。它的第一个版本包括3.56亿个查询、1.66亿个搜索结果页面和55家搜索提供商提供的17亿个搜索结果。尽管已经在文献中研究了许多查询日志，但拥有这些数据的搜索提供商通常不会公开日志以保护用户隐私和重要的业务数据。在少数公开可用的查询日志中，没有一个结合了规模、范围和多样性。AQL是第一个这样做的日志，可以进行新的检索模型和（历时）搜索引擎分析的研究。以一种保护隐私的方式提供，它促进了开放式研究以及搜索行业更多的透明度和问责制。

0

相关内容

查询日志

【元宇宙】“The State Of The Metaverse”26页报告

【元宇宙】“The State Of The Metaverse”26页报告

专知会员服务

45+阅读 · 2022年5月25日

Artificial Intelligence: Ready to Ride the Wave? BCG 28页PPT

Artificial Intelligence: Ready to Ride the Wave? BCG 28页PPT

专知会员服务

28+阅读 · 2022年2月20日

【SIGIR2020】学习搜索查询的颜色表示，Learning Colour Representations of Search Queries

【SIGIR2020】学习搜索查询的颜色表示，Learning Colour Representations of Search Queries

专知会员服务

17+阅读 · 2020年6月18日

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

专知会员服务

33+阅读 · 2020年5月2日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

SIGIR2019 接收论文列表

SIGIR2019 接收论文列表

专知

18+阅读 · 2019年4月20日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇知识图谱相关论文—嵌入式知识、Zero-shot识别、知识图谱嵌入、网络库、变分推理、解释、弱监督

【论文推荐】最新七篇知识图谱相关论文—嵌入式知识、Zero-shot识别、知识图谱嵌入、网络库、变分推理、解释、弱监督

专知

19+阅读 · 2018年3月26日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

Zakharov系统的解的动力学行为研究

国家自然科学基金

0+阅读 · 2015年12月31日

hsa-miR-129-5p负调节Warburg效应和抑制肝癌生长转移的分子机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

生物炭与土壤粘粒矿物及其复合体的耦合效应

国家自然科学基金

0+阅读 · 2013年12月31日

AM真菌对锑在土壤-植物系统中迁移转化的作用机理

国家自然科学基金

0+阅读 · 2013年12月31日

夏闲期日光温室土壤氧化亚氮排放及驱动机制

国家自然科学基金

0+阅读 · 2012年12月31日

番茄抗病膜蛋白TARK1稳定性的调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

拟南芥泛素连接酶BAR1在植物天然免疫中的功能

国家自然科学基金

0+阅读 · 2012年12月31日

模块化非线性系统辨识

国家自然科学基金

0+阅读 · 2011年12月31日

污染土壤中磷肥影响铜植物有效性机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

Multimodal User Authentication in Smart Environments: Survey of User Attitudes

Arxiv

0+阅读 · 2023年5月23日

Knowledge Refinement via Interaction Between Search Engines and Large Language Models

Arxiv

0+阅读 · 2023年5月21日

On recoverability from failures in dual voting

Arxiv

0+阅读 · 2023年5月20日

An Experimental Investigation of Tuning QUIC-Based Publish-Subscribe Architectures in IoT

Arxiv

0+阅读 · 2023年5月19日

Summarizing Strategy Card Game AI Competition

Arxiv

0+阅读 · 2023年5月19日

Trustworthy, responsible, ethical AI in manufacturing and supply chains: synthesis and emerging research questions

Arxiv

0+阅读 · 2023年5月19日

Semi-verified PAC Learning from the Crowd

Arxiv

0+阅读 · 2023年5月18日

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Arxiv

16+阅读 · 2022年11月3日

Artificial Intelligence for the Metaverse: A Survey

Arxiv

31+阅读 · 2022年2月15日

Mining Disinformation and Fake News: Concepts, Methods, and Recent Advancements

Mining Disinformation and Fake News: Concepts, Methods, and Recent Advancements

Arxiv

16+阅读 · 2020年1月2日

VIP会员

文章信息

相关主题

相关VIP内容

【元宇宙】“The State Of The Metaverse”26页报告

【元宇宙】“The State Of The Metaverse”26页报告

专知会员服务

45+阅读 · 2022年5月25日

Artificial Intelligence: Ready to Ride the Wave? BCG 28页PPT

Artificial Intelligence: Ready to Ride the Wave? BCG 28页PPT

专知会员服务

28+阅读 · 2022年2月20日

【SIGIR2020】学习搜索查询的颜色表示，Learning Colour Representations of Search Queries

【SIGIR2020】学习搜索查询的颜色表示，Learning Colour Representations of Search Queries

专知会员服务

17+阅读 · 2020年6月18日

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

【2020关键词提取】医学报告的关键词提取和结构化，Keyword extraction and structuralization of medical reports

专知会员服务

33+阅读 · 2020年5月2日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

【人工智能在2019：一年回顾】反人工智能，AI in 2019: A Year in Review

专知会员服务

79+阅读 · 2019年10月10日

【哈佛大学商学院课程Fall 2019】机器学习可解释性

【哈佛大学商学院课程Fall 2019】机器学习可解释性

专知会员服务

105+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

小规模训练指南：打造世界级大语言模型的关键方法

无人机编队飞行：复杂环境中作战的策略、挑战与应用

大模型APP，AI时代第一个爆款

从数据中心视角出发的高效大语言模型训练综述

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

SIGIR2019 接收论文列表

SIGIR2019 接收论文列表

专知

18+阅读 · 2019年4月20日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇知识图谱相关论文—嵌入式知识、Zero-shot识别、知识图谱嵌入、网络库、变分推理、解释、弱监督

【论文推荐】最新七篇知识图谱相关论文—嵌入式知识、Zero-shot识别、知识图谱嵌入、网络库、变分推理、解释、弱监督

专知

19+阅读 · 2018年3月26日

【推荐】用Python/OpenCV实现增强现实

【推荐】用Python/OpenCV实现增强现实

机器学习研究会

15+阅读 · 2017年11月16日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

Multimodal User Authentication in Smart Environments: Survey of User Attitudes

Arxiv

0+阅读 · 2023年5月23日

Knowledge Refinement via Interaction Between Search Engines and Large Language Models

Arxiv

0+阅读 · 2023年5月21日

On recoverability from failures in dual voting

Arxiv

0+阅读 · 2023年5月20日

An Experimental Investigation of Tuning QUIC-Based Publish-Subscribe Architectures in IoT

Arxiv

0+阅读 · 2023年5月19日

Summarizing Strategy Card Game AI Competition

Arxiv

0+阅读 · 2023年5月19日

Trustworthy, responsible, ethical AI in manufacturing and supply chains: synthesis and emerging research questions

Arxiv

0+阅读 · 2023年5月19日

Semi-verified PAC Learning from the Crowd

Arxiv

0+阅读 · 2023年5月18日

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Arxiv

16+阅读 · 2022年11月3日

Artificial Intelligence for the Metaverse: A Survey

Arxiv

31+阅读 · 2022年2月15日

Mining Disinformation and Fake News: Concepts, Methods, and Recent Advancements

Mining Disinformation and Fake News: Concepts, Methods, and Recent Advancements

Arxiv

16+阅读 · 2020年1月2日

相关基金

Zakharov系统的解的动力学行为研究

国家自然科学基金

0+阅读 · 2015年12月31日

hsa-miR-129-5p负调节Warburg效应和抑制肝癌生长转移的分子机理研究

国家自然科学基金

0+阅读 · 2014年12月31日

长链非编码RNA CAR intergenic 10在细胞衰老中的作用和机制

国家自然科学基金

1+阅读 · 2013年12月31日

生物炭与土壤粘粒矿物及其复合体的耦合效应

国家自然科学基金

0+阅读 · 2013年12月31日

AM真菌对锑在土壤-植物系统中迁移转化的作用机理

国家自然科学基金

0+阅读 · 2013年12月31日

夏闲期日光温室土壤氧化亚氮排放及驱动机制

国家自然科学基金

0+阅读 · 2012年12月31日

番茄抗病膜蛋白TARK1稳定性的调控机理研究

国家自然科学基金

0+阅读 · 2012年12月31日

拟南芥泛素连接酶BAR1在植物天然免疫中的功能

国家自然科学基金

0+阅读 · 2012年12月31日

模块化非线性系统辨识

国家自然科学基金

0+阅读 · 2011年12月31日

污染土壤中磷肥影响铜植物有效性机制研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员