Unfolded Self-Reconstruction LSH: 面向近似最近邻搜索的机器遗忘 (Unfolded Self-Reconstruction LSH: Towards Machine Unlearning in Approximate Nearest Neighbour Search) - 专知论文

会员服务 ·

0

LSH · 哈希 · 最近邻 · 近邻 · 哈希方法 ·

2023 年 4 月 5 日

Unfolded Self-Reconstruction LSH: Towards Machine Unlearning in Approximate Nearest Neighbour Search

翻译：Unfolded Self-Reconstruction LSH: 面向近似最近邻搜索的机器遗忘

Kim Yong Tan,Lyu Yueming,Yew-Soon Ong,Ivor Tsang

Approximate nearest neighbour (ANN) search is an essential component of search engines, recommendation systems, etc. Many recent works focus on learning-based data-distribution-dependent hashing and achieve good retrieval performance. However, due to increasing demand for users' privacy and security, we often need to remove users' data information from Machine Learning (ML) models to satisfy specific privacy and security requirements. This need requires the ANN search algorithm to support fast online data deletion and insertion. Current learning-based hashing methods need retraining the hash function, which is prohibitable due to the vast time-cost of large-scale data. To address this problem, we propose a novel data-dependent hashing method named unfolded self-reconstruction locality-sensitive hashing (USR-LSH). Our USR-LSH unfolded the optimization update for instance-wise data reconstruction, which is better for preserving data information than data-independent LSH. Moreover, our USR-LSH supports fast online data deletion and insertion without retraining. To the best of our knowledge, we are the first to address the machine unlearning of retrieval problems. Empirically, we demonstrate that USR-LSH outperforms the state-of-the-art data-distribution-independent LSH in ANN tasks in terms of precision and recall. We also show that USR-LSH has significantly faster data deletion and insertion time than learning-based data-dependent hashing.

翻译：近似最近邻（ANN）搜索是搜索引擎、推荐系统等的重要组成部分。许多最近的工作关注于基于学习的数据分布依赖哈希，取得了良好的检索性能。然而，由于用户隐私和安全的需求不断增加，我们常常需要从机器学习（ML）模型中删除用户数据信息，以满足特定的隐私和安全要求。这种需求需要ANN搜索算法支持快速的在线数据删除和插入。目前的基于学习的哈希方法需要重新训练哈希函数，由于大规模数据的巨大时间成本，这是不可接受的。为了解决这个问题，我们提出了一种新的数据依赖性哈希方法，称为unfolded self-reconstruction locality-sensitive hashing（USR-LSH）。我们的USR-LSH展开了基于实例的数据重构的优化更新，这比基于数据相互独立的LSH更好地保留了数据信息。此外，我们的USR-LSH支持快速的在线数据删除和插入，无需重新训练。据我们所知，我们是第一个解决检索问题机器遗忘的人。在实证方面，我们证明了USR-LSH在ANN任务中的检索性能（精度和召回率）优于现有的基于数据分布相互独立的LSH方法。我们还表明USR-LSH比基于学习的数据依赖哈希具有显著更快的数据删除和插入时间。

0

相关内容

LSH

局部敏感哈希算法

【CVPR2023】正则化二阶影响的持续学习

【CVPR2023】正则化二阶影响的持续学习

专知会员服务

19+阅读 · 2023年4月22日

WWW21最新「比较学习」教程，135页PPT阐述从排名数据中学习

专知会员服务

37+阅读 · 2021年4月27日

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

专知会员服务

74+阅读 · 2020年7月6日

【KDD2020】基于矩阵和张量因子分解的高效自动机器学习搜索，Efficient AutoML Pipeline Search with Matrix and Tensor Factorization

【KDD2020】基于矩阵和张量因子分解的高效自动机器学习搜索，Efficient AutoML Pipeline Search with Matrix and Tensor Factorization

专知会员服务

13+阅读 · 2020年6月10日

【SIGMOD2020】一个全面的主动学习方法的实体匹配基准框架，A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching

【SIGMOD2020】一个全面的主动学习方法的实体匹配基准框架，A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching

专知会员服务

24+阅读 · 2020年3月31日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【Uber AI新论文】持续元学习，Learning to Continually Learn

【Uber AI新论文】持续元学习，Learning to Continually Learn

专知会员服务

37+阅读 · 2020年2月27日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

【康奈尔大学】度量数据粒度，Measuring Dataset Granularity

【康奈尔大学】度量数据粒度，Measuring Dataset Granularity

专知会员服务

13+阅读 · 2019年12月27日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

模型会忘了你是谁吗？两篇Machine Unlearning顶会论文告诉你什么是模型遗忘

模型会忘了你是谁吗？两篇Machine Unlearning顶会论文告诉你什么是模型遗忘

PaperWeekly

5+阅读 · 2022年10月9日

WWW2022 | Recommendation Unlearning

WWW2022 | Recommendation Unlearning

机器学习与推荐算法

0+阅读 · 2022年6月2日

【Uber AI新论文】持续元学习，Learning to Continually Learn

【Uber AI新论文】持续元学习，Learning to Continually Learn

专知

19+阅读 · 2020年2月27日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

面向在线检索的医学影像多特征降维方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向大数据跨媒体检索的多模态哈希学习方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

适定的多元样条逼近方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

面向异构环境的多任务多视图学习算法研究

国家自然科学基金

3+阅读 · 2014年12月31日

低占空比无线传感器网络跨层协同传输关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

有理函数非旋转Fatou域与不连通Julia集的结构

国家自然科学基金

0+阅读 · 2014年12月31日

海量、动态、嘈杂语义数据集上的递增随时推理方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

动态云环境中基于SLA的工作流调度

国家自然科学基金

0+阅读 · 2012年12月31日

面向视频目标识别的图像集合分类方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

He和H离子注入Si基材料引起的表面剥离及机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

Constrained Proximal Policy Optimization

Arxiv

0+阅读 · 2023年5月23日

Term-Sets Can Be Strong Document Identifiers For Auto-Regressive Search Engines

Arxiv

0+阅读 · 2023年5月23日

Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline

Arxiv

0+阅读 · 2023年5月22日

Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models

Arxiv

1+阅读 · 2023年5月22日

Towards Robust and Accurate Myoelectric Controller Design based on Multi-objective Optimization using Evolutionary Computation

Arxiv

0+阅读 · 2023年5月22日

ConQueR: Contextualized Query Reduction using Search Logs

Arxiv

0+阅读 · 2023年5月22日

Self-Reinforcement Attention Mechanism For Tabular Learning

Arxiv

0+阅读 · 2023年5月19日

Evolutionary Diversity Optimisation in Constructing Satisfying Assignments

Arxiv

0+阅读 · 2023年5月19日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

VIP会员

文章信息

相关主题

相关VIP内容

【CVPR2023】正则化二阶影响的持续学习

【CVPR2023】正则化二阶影响的持续学习

专知会员服务

19+阅读 · 2023年4月22日

WWW21最新「比较学习」教程，135页PPT阐述从排名数据中学习

专知会员服务

37+阅读 · 2021年4月27日

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

专知会员服务

74+阅读 · 2020年7月6日

【KDD2020】基于矩阵和张量因子分解的高效自动机器学习搜索，Efficient AutoML Pipeline Search with Matrix and Tensor Factorization

【KDD2020】基于矩阵和张量因子分解的高效自动机器学习搜索，Efficient AutoML Pipeline Search with Matrix and Tensor Factorization

专知会员服务

13+阅读 · 2020年6月10日

【SIGMOD2020】一个全面的主动学习方法的实体匹配基准框架，A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching

【SIGMOD2020】一个全面的主动学习方法的实体匹配基准框架，A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching

专知会员服务

24+阅读 · 2020年3月31日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【Uber AI新论文】持续元学习，Learning to Continually Learn

【Uber AI新论文】持续元学习，Learning to Continually Learn

专知会员服务

37+阅读 · 2020年2月27日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

【康奈尔大学】度量数据粒度，Measuring Dataset Granularity

【康奈尔大学】度量数据粒度，Measuring Dataset Granularity

专知会员服务

13+阅读 · 2019年12月27日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

《乌克兰无人机产业：志愿者与政策在构建新兴无人机产业中的协同作用》最新报告

《人工智能辅助决策中的数据可视化：系统性综述》

人工智能驱动弹药制造现代化：美国陆军转型之路

《敏捷作战部署中枢纽-辐条基地选址优化研究》80页

相关资讯

模型会忘了你是谁吗？两篇Machine Unlearning顶会论文告诉你什么是模型遗忘

模型会忘了你是谁吗？两篇Machine Unlearning顶会论文告诉你什么是模型遗忘

PaperWeekly

5+阅读 · 2022年10月9日

WWW2022 | Recommendation Unlearning

WWW2022 | Recommendation Unlearning

机器学习与推荐算法

0+阅读 · 2022年6月2日

【Uber AI新论文】持续元学习，Learning to Continually Learn

【Uber AI新论文】持续元学习，Learning to Continually Learn

专知

19+阅读 · 2020年2月27日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

基于PyTorch/TorchText的自然语言处理库

基于PyTorch/TorchText的自然语言处理库

专知

28+阅读 · 2019年4月22日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

相关论文

Constrained Proximal Policy Optimization

Arxiv

0+阅读 · 2023年5月23日

Term-Sets Can Be Strong Document Identifiers For Auto-Regressive Search Engines

Arxiv

0+阅读 · 2023年5月23日

Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline

Arxiv

0+阅读 · 2023年5月22日

Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models

Arxiv

1+阅读 · 2023年5月22日

Towards Robust and Accurate Myoelectric Controller Design based on Multi-objective Optimization using Evolutionary Computation

Arxiv

0+阅读 · 2023年5月22日

ConQueR: Contextualized Query Reduction using Search Logs

Arxiv

0+阅读 · 2023年5月22日

Self-Reinforcement Attention Mechanism For Tabular Learning

Arxiv

0+阅读 · 2023年5月19日

Evolutionary Diversity Optimisation in Constructing Satisfying Assignments

Arxiv

0+阅读 · 2023年5月19日

Introduction to Online Convex Optimization

Arxiv

23+阅读 · 2021年12月19日

Self-correcting Q-Learning

Arxiv

11+阅读 · 2020年12月2日

相关基金

面向在线检索的医学影像多特征降维方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

面向大数据跨媒体检索的多模态哈希学习方法研究

国家自然科学基金

0+阅读 · 2015年12月31日

适定的多元样条逼近方法研究

国家自然科学基金

0+阅读 · 2014年12月31日

面向异构环境的多任务多视图学习算法研究

国家自然科学基金

3+阅读 · 2014年12月31日

低占空比无线传感器网络跨层协同传输关键技术研究

国家自然科学基金

0+阅读 · 2014年12月31日

有理函数非旋转Fatou域与不连通Julia集的结构

国家自然科学基金

0+阅读 · 2014年12月31日

海量、动态、嘈杂语义数据集上的递增随时推理方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

动态云环境中基于SLA的工作流调度

国家自然科学基金

0+阅读 · 2012年12月31日

面向视频目标识别的图像集合分类方法研究

国家自然科学基金

0+阅读 · 2012年12月31日

He和H离子注入Si基材料引起的表面剥离及机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员