改进激光组以获取高维绝对数据 (Improving Group Lasso for high-dimensional categorical data) - 专知论文

会员服务 ·

0

Group Lasso · 分类数据 · GROUP · MoDELS · 稀疏 ·

2022 年 10 月 27 日

Improving Group Lasso for high-dimensional categorical data

翻译：改进激光组以获取高维绝对数据

Szymon Nowakowski,Piotr Pokarowski,Wojciech Rejchel

from arxiv, arXiv admin note: text overlap with arXiv:2112.11114

Sparse modelling or model selection with categorical data is challenging even for a moderate number of variables, because one parameter is roughly needed to encode one category or level. The Group Lasso is a well known efficient algorithm for selection continuous or categorical variables, but all estimates related to a selected factor usually differ. Therefore, a fitted model may not be sparse, which makes the model interpretation difficult. To obtain a sparse solution of the Group Lasso we propose the following two-step procedure: first, we reduce data dimensionality using the Group Lasso; then to choose the final model we use an information criterion on a small family of models prepared by clustering levels of individual factors. We investigate selection correctness of the algorithm in a sparse high-dimensional scenario. We also test our method on synthetic as well as real datasets and show that it performs better than the state of the art algorithms with respect to the prediction accuracy or model dimension.

翻译：即使对于数量不多的变数来说,使用绝对数据进行粗略的建模或模型选择也具有挑战性,因为对于一个类别或层次的编码,大致需要有一个参数。Lasso集团是一个众所周知的用于选择连续或绝对变量的有效算法,但所有与选定因素有关的估计通常各不相同。因此,一个合适的模型可能并不稀疏,因此模型解释难于使用。为了获得Lasso集团的稀疏解决方案,我们建议采用以下两步程序:首先,我们使用Lasso集团来减少数据维度;然后选择我们使用的信息标准来选择一个最后模型,我们使用由个别因素组合层次所制作的模型组成的小系列信息标准。我们调查在稀疏高维情景中选择算法的正确性。我们还在合成和真实数据集方面测试我们的方法,并显示它比预测准确性或模型维度的先进算法状态要好。

0

相关内容

Group Lasso

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【斯坦福大学博士论文】大规模和高维统计学习方法和算法，147页pdf， Large-scale and high-dimensional statistical learning methods and algorithms

专知会员服务

26+阅读 · 2020年6月13日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

基于张量稀疏L1图的半监督极化SAR影像地物分类

国家自然科学基金

0+阅读 · 2015年12月31日

植物分子设计中高维数据的低维稀疏逼近方法

国家自然科学基金

0+阅读 · 2015年12月31日

状态空间搜索的anytime模式及其高效算法研究

国家自然科学基金

0+阅读 · 2015年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

协同主被动光学遥感数据的多尺度森林叶面积指数反演研究

国家自然科学基金

0+阅读 · 2013年12月31日

空间相依数据的统计推断及其应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

蓖麻Geranylgeranyl Reductase酶在维生素E高效合成途径中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

因果推断的统计方法

国家自然科学基金

26+阅读 · 2011年12月31日

矩阵分解的低延迟并行算法

国家自然科学基金

0+阅读 · 2009年12月31日

一维动力系统的Julia集及其不变子集的维数与熵

国家自然科学基金

0+阅读 · 2009年12月31日

Flexible Regularized Estimation in High-Dimensional Mixed Membership Models

Arxiv

0+阅读 · 2022年12月13日

Towards Efficient and Domain-Agnostic Evasion Attack with High-dimensional Categorical Inputs

Arxiv

0+阅读 · 2022年12月13日

Dimensionality reduction on complex vector spaces for dynamic weighted Euclidean distance

Arxiv

0+阅读 · 2022年12月13日

Tractability of $L_2$-approximation and integration in weighted Hermite spaces of finite smoothness

Arxiv

0+阅读 · 2022年12月12日

Weak signal identification and inference in penalized likelihood models for categorical responses

Arxiv

0+阅读 · 2022年12月12日

Bivariate Causal Discovery for Categorical Data via Classification with Optimal Label Permutation

Arxiv

0+阅读 · 2022年12月11日

Polynomial Distributions and Transformations

Arxiv

0+阅读 · 2022年12月9日

Direct sampling method to inverse wave-number-dependent source problems (part I): determination of the support of a stationary source

Arxiv

0+阅读 · 2022年12月9日

Model-based clustering of categorical data based on the Hamming distance

Arxiv

0+阅读 · 2022年12月9日

Modern Statistical Models and Methods for Estimating Fatigue-Life and Fatigue-Strength Distributions from Experimental Data

Arxiv

0+阅读 · 2022年12月8日

VIP会员

文章信息

相关主题

相关VIP内容

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

75+阅读 · 2022年6月28日

Linux导论，Introduction to Linux，96页ppt

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

【斯坦福大学博士论文】大规模和高维统计学习方法和算法，147页pdf， Large-scale and high-dimensional statistical learning methods and algorithms

专知会员服务

26+阅读 · 2020年6月13日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

2019年机器学习框架回顾

2019年机器学习框架回顾

专知会员服务

36+阅读 · 2019年10月11日

机器学习入门的经验与建议

机器学习入门的经验与建议

专知会员服务

94+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

【新书】面向企业的图学习扩展：生产级图学习与推理，485页pdf

AI智能体编程：技术、挑战与机遇综述

【国家标准】数据安全技术数据安全风险评估方法

【CMU博士论文】交互式学习的进展：替代性反馈机制与自适应因果推理

相关资讯

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Latest News & Announcements of the Workshop

【ICIG2021】Latest News & Announcements of the Workshop

中国图象图形学学会CSIG

0+阅读 · 2021年12月20日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium1

中国图象图形学学会CSIG

0+阅读 · 2021年11月3日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

相关论文

Flexible Regularized Estimation in High-Dimensional Mixed Membership Models

Arxiv

0+阅读 · 2022年12月13日

Towards Efficient and Domain-Agnostic Evasion Attack with High-dimensional Categorical Inputs

Arxiv

0+阅读 · 2022年12月13日

Dimensionality reduction on complex vector spaces for dynamic weighted Euclidean distance

Arxiv

0+阅读 · 2022年12月13日

Tractability of $L_2$-approximation and integration in weighted Hermite spaces of finite smoothness

Arxiv

0+阅读 · 2022年12月12日

Weak signal identification and inference in penalized likelihood models for categorical responses

Arxiv

0+阅读 · 2022年12月12日

Bivariate Causal Discovery for Categorical Data via Classification with Optimal Label Permutation

Arxiv

0+阅读 · 2022年12月11日

Polynomial Distributions and Transformations

Arxiv

0+阅读 · 2022年12月9日

Direct sampling method to inverse wave-number-dependent source problems (part I): determination of the support of a stationary source

Arxiv

0+阅读 · 2022年12月9日

Model-based clustering of categorical data based on the Hamming distance

Arxiv

0+阅读 · 2022年12月9日

Modern Statistical Models and Methods for Estimating Fatigue-Life and Fatigue-Strength Distributions from Experimental Data

Arxiv

0+阅读 · 2022年12月8日

相关基金

基于张量稀疏L1图的半监督极化SAR影像地物分类

国家自然科学基金

0+阅读 · 2015年12月31日

植物分子设计中高维数据的低维稀疏逼近方法

国家自然科学基金

0+阅读 · 2015年12月31日

状态空间搜索的anytime模式及其高效算法研究

国家自然科学基金

0+阅读 · 2015年12月31日

采用pinball loss的MEE算法研究

国家自然科学基金

1+阅读 · 2013年12月31日

协同主被动光学遥感数据的多尺度森林叶面积指数反演研究

国家自然科学基金

0+阅读 · 2013年12月31日

空间相依数据的统计推断及其应用研究

国家自然科学基金

0+阅读 · 2012年12月31日

蓖麻Geranylgeranyl Reductase酶在维生素E高效合成途径中的作用研究

国家自然科学基金

0+阅读 · 2012年12月31日

因果推断的统计方法

国家自然科学基金

26+阅读 · 2011年12月31日

矩阵分解的低延迟并行算法

国家自然科学基金

0+阅读 · 2009年12月31日

一维动力系统的Julia集及其不变子集的维数与熵

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员