以k美元为单位的新约束值和信息以k美元为单位 (New bounds for $k$-means and information $k$-means) - 专知论文

会员服务 ·

0

INFORMS · 泛化理论 · 核技巧 · 准则 · 样本 ·

2021 年 1 月 14 日

New bounds for $k$-means and information $k$-means

翻译：以k美元为单位的新约束值和信息以k美元为单位

Gautier Appert,Olivier Catoni

In this paper, we derive a new dimension-free non-asymptotic upper bound for the quadratic $k$-means excess risk related to the quantization of an i.i.d sample in a separable Hilbert space. We improve the bound of order $\mathcal{O} \bigl( k / \sqrt{n} \bigr)$ of Biau, Devroye and Lugosi, by establishing a bound of order $\mathcal{O} \bigl(\log(n/k) \sqrt{k \log(k) / n} \, \bigr)$ where $k$ is the number of centers and $n$ the sample size. This is essentially optimal up to logarithmic factors since a lower bound of order $\mathcal{O} \bigl( \sqrt{k^{1 - 4/d}/n} \bigr)$ is known in dimension $d$. Our technique of proof is based on the linearization of the $k$-means criterion through a kernel trick and on PAC-Bayesian inequalities. To get a $1 / \sqrt{n}$ speed, we introduce a new PAC-Bayesian chaining method replacing the concept of $\delta$-net with the perturbation of the parameter by an infinite dimensional Gaussian process. In the meantime, we embed the usual $k$-means criterion into a broader family built upon the Kullback divergence and its underlying properties. This results in a new algorithm that we named information $k$-means, well suited to the clustering of bags of words. Based on considerations from information theory, we also introduce a new bounded $k$-means criterion that uses a scale parameter but satisfies a generalization bound that does not require any boundedness or even integrability conditions on the sample. We describe the counterpart of Lloyd's algorithm and prove generalization bounds for these new $k$-means criteria.

翻译：在本文中, 我们得出一个新的无维度的上方基流, 用于四维值 $k$- 表示美元( miglation) 。在可分解的 Hilbert 空间中, i. d 样本的四分位化超风险。我们改进了 $\ mathcal{ O}\ bigl( k/\ sqrt{ n}\ bigr) 美元( bigl) 的上方基值。通过建立 $( sqrt{ mathal{O}\ biglocklation) 的组合, 美元( more) 美元( more) 美元( liver) 代表美元( lider) 美元( lider) 。美元( liver) =( liver) =( liver) 值( liver) =( liver) =( liver) =( lax) a legal lax) a likeal legal legal modeal) a a le le lemental cal lex) a a modeal modeal modeal cal a.

0

相关内容

INFORMS

《计算机信息》杂志发表高质量的论文，扩大了运筹学和计算的范围，寻求有关理论、方法、实验、系统和应用方面的原创研究论文、新颖的调查和教程论文，以及描述新的和有用的软件工具的论文。官网链接：https://pubsonline.informs.org/journal/ijoc

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

54+阅读 · 2020年9月7日

一份简单《图神经网络》教程，28页ppt

一份简单《图神经网络》教程，28页ppt

专知会员服务

127+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

已删除

将门创投

3+阅读 · 2019年4月12日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

Ray RLlib: Scalable 降龙十八掌

Ray RLlib: Scalable 降龙十八掌

CreateAMind

9+阅读 · 2018年12月28日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【NIPS2018】接收论文列表

【NIPS2018】接收论文列表

专知

5+阅读 · 2018年9月10日

【论文推荐】最新六篇主题模型相关论文—领域特定知识库、神经变分推断、动态和静态主题模型

【论文推荐】最新六篇主题模型相关论文—领域特定知识库、神经变分推断、动态和静态主题模型

专知

19+阅读 · 2018年6月26日

Hierarchical Disentangled Representations

Hierarchical Disentangled Representations

CreateAMind

4+阅读 · 2018年4月15日

【学习】Hierarchical Softmax

【学习】Hierarchical Softmax

机器学习研究会

4+阅读 · 2017年8月6日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

Bounds on half graph orders in powers of sparse graphs

Bounds on half graph orders in powers of sparse graphs

Arxiv

0+阅读 · 2021年3月10日

On Information Gain and Regret Bounds in Gaussian Process Bandits

Arxiv

0+阅读 · 2021年3月9日

Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds

Arxiv

0+阅读 · 2021年3月9日

Improved upper bounds for the rigidity of Kronecker products

Arxiv

0+阅读 · 2021年3月9日

Structure Assisted NMF Methods for Separation of Degenerate Mixture Data with Application to NMR Spectroscopy

Arxiv

0+阅读 · 2021年3月9日

A Lower Bound for the Sample Complexity of Inverse Reinforcement Learning

Arxiv

0+阅读 · 2021年3月7日

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

Arxiv

0+阅读 · 2021年3月7日

Numerical results for an unconditionally stable space-time finite element method for the wave equation

Arxiv

0+阅读 · 2021年3月7日

New Separations Results for External Information

Arxiv

0+阅读 · 2021年3月6日

Variational Bayesian Reinforcement Learning with Regret Bounds

Arxiv

3+阅读 · 2018年7月25日

VIP会员

文章信息

相关主题

相关VIP内容

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

54+阅读 · 2020年9月7日

一份简单《图神经网络》教程，28页ppt

一份简单《图神经网络》教程，28页ppt

专知会员服务

127+阅读 · 2020年8月2日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

热门VIP内容

开通专知VIP会员享更多权益服务

《城市滨海地区：理解复杂多变环境下的指挥控制框架》50页报告

《理解城市战及其在俄乌战争中的表现》报告

美空军“顶点2025”实验：推进AI在C2、动态目标锁定与联盟集成中的应用

《建设式兵棋模拟作为战术集群配置优化的关键组成部分》

相关资讯

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

已删除

将门创投

3+阅读 · 2019年4月12日

RL 真经

CreateAMind

5+阅读 · 2018年12月28日

Ray RLlib: Scalable 降龙十八掌

Ray RLlib: Scalable 降龙十八掌

CreateAMind

9+阅读 · 2018年12月28日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【NIPS2018】接收论文列表

【NIPS2018】接收论文列表

专知

5+阅读 · 2018年9月10日

【论文推荐】最新六篇主题模型相关论文—领域特定知识库、神经变分推断、动态和静态主题模型

【论文推荐】最新六篇主题模型相关论文—领域特定知识库、神经变分推断、动态和静态主题模型

专知

19+阅读 · 2018年6月26日

Hierarchical Disentangled Representations

Hierarchical Disentangled Representations

CreateAMind

4+阅读 · 2018年4月15日

【学习】Hierarchical Softmax

【学习】Hierarchical Softmax

机器学习研究会

4+阅读 · 2017年8月6日

Auto-Encoding GAN

Auto-Encoding GAN

CreateAMind

7+阅读 · 2017年8月4日

相关论文

Bounds on half graph orders in powers of sparse graphs

Bounds on half graph orders in powers of sparse graphs

Arxiv

0+阅读 · 2021年3月10日

On Information Gain and Regret Bounds in Gaussian Process Bandits

Arxiv

0+阅读 · 2021年3月9日

Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds

Arxiv

0+阅读 · 2021年3月9日

Improved upper bounds for the rigidity of Kronecker products

Arxiv

0+阅读 · 2021年3月9日

Structure Assisted NMF Methods for Separation of Degenerate Mixture Data with Application to NMR Spectroscopy

Arxiv

0+阅读 · 2021年3月9日

A Lower Bound for the Sample Complexity of Inverse Reinforcement Learning

Arxiv

0+阅读 · 2021年3月7日

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

Arxiv

0+阅读 · 2021年3月7日

Numerical results for an unconditionally stable space-time finite element method for the wave equation

Arxiv

0+阅读 · 2021年3月7日

New Separations Results for External Information

Arxiv

0+阅读 · 2021年3月6日

Variational Bayesian Reinforcement Learning with Regret Bounds

Arxiv

3+阅读 · 2018年7月25日

微信扫码咨询专知VIP会员