多任务回归预测的样本大小依赖性通用共变数子集的最佳选择 (Optimal selection of sample-size dependent common subsets of covariates for multi-task regression prediction) - 专知论文

会员服务 ·

0

训练集 · 情景 · 数据集 · 优化器 · 随机采样 ·

2021 年 9 月 5 日

Optimal selection of sample-size dependent common subsets of covariates for multi-task regression prediction

翻译：多任务回归预测的样本大小依赖性通用共变数子集的最佳选择

David Azriel,Yosef Rinott

An analyst is given a training set consisting of regression datasets $D_j$ of different sizes, which are distributed according to some $G_j$, $j=1,\ldots,\cal J$, where the distributions $G_j$ are assumed to form a random sample generated by some common source. In particular, the $D_j$'s have a common set of covariates and they are all labeled. The training set is used by the analyst for selection of subsets of covariates denoted by ${P}^*(n)$, whose role is described next. The multi-task problem we consider is as follows: given a number of random labeled datasets (which may be in the training set or not) $D_{J_k}$ of size $n_k$, $k=1,\ldots,K$, estimate separately for each dataset the regression coefficients on the subset of covariates ${P}^*(n_k)$ and then predict future dependent variables given their covariates. Naturally, a large sample size $n_k$ of $D_{J_k}$ allows a larger subset of covariates, and the dependence of the size of the selected covariate subsets on $n_k$ is needed in order to achieve good prediction and avoid overfitting. Subset selection is notoriously difficult and computationally demanding, and requires large samples; using all the regression datasets in the training set together amounts to borrowing strength toward better selection under suitable assumptions. Furthermore, using common subsets for all regressions having a given sample size standardizes and simplifies the data collection and avoids having to select and use a different subset for each prediction task. Our approach is efficient when the relevant covariates for prediction are common to the different regressions, while the models' coefficients may vary between different regressions.

翻译：向分析师提供一套由回归数据集组成的培训组, 该组由不同大小的回归数据集组成, 这些数据集根据一些 G_ j$, $j=1,\ldots,\cal J$, 其中分配 $G_ j$假设形成由某些共同来源产生的随机抽样。特别是, $D_ j$有一套共同的共变数, 它们都有标签。分析师使用这套培训组来选择由 ${P% (n) =(n) 美元表示的共变数子子子数, 其作用将在下文加以描述。我们考虑的多任务组数问题如下: 随机标定的数据集数( 可能在训练组中出现 ) $_ j_ j$, 以随机随机抽样样本数, 美元, 美元=1,\ldots, K美元, 分别估算每组数据在计算 $ (n_) (n_k) rick$) 的子数时, 将计算回归系数系数, 然后预测未来依次数预测。当然, 需要大美元和 road_ codeal_ deal codeal 。

0

相关内容

训练集

训练集，在AI领域多指用于机器学习训练的数据，数据可以有标签的，也可以是无标签的。

【经典书】线性代数，436页pdf

专知会员服务

78+阅读 · 2021年3月16日

【经典书】线性代数元素，197页pdf

【经典书】线性代数元素，197页pdf

专知会员服务

57+阅读 · 2021年3月4日

最新《图理论》笔记书，98页pdf

最新《图理论》笔记书，98页pdf

专知会员服务

76+阅读 · 2020年12月27日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

【Manning新书】现代Java实战，592页pdf

【Manning新书】现代Java实战，592页pdf

专知会员服务

101+阅读 · 2020年5月22日

【ICLR2020】理解非自回归机器翻译中的知识蒸馏（Understanding Knowledge Distillation in Non-autoregressive Machine Translation）

【ICLR2020】理解非自回归机器翻译中的知识蒸馏（Understanding Knowledge Distillation in Non-autoregressive Machine Translation）

专知会员服务

11+阅读 · 2019年12月28日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

【论文笔记】通俗理解少样本文本分类 (Few-Shot Text Classification) (1)

【论文笔记】通俗理解少样本文本分类 (Few-Shot Text Classification) (1)

深度学习自然语言处理

7+阅读 · 2020年4月8日

分布式并行架构Ray介绍

分布式并行架构Ray介绍

CreateAMind

10+阅读 · 2019年8月9日

已删除

将门创投

5+阅读 · 2019年6月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

分布式TensorFlow入门指南

分布式TensorFlow入门指南

机器学习研究会

4+阅读 · 2017年11月28日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【学习】(Python)SVM数据分类

【学习】(Python)SVM数据分类

机器学习研究会

6+阅读 · 2017年10月15日

【推荐】决策树/随机森林深入解析

【推荐】决策树/随机森林深入解析

机器学习研究会

5+阅读 · 2017年9月21日

Inference in Regression Discontinuity Designs with High-Dimensional Covariates

Arxiv

0+阅读 · 2021年10月26日

Sample Selection Bias in Evaluation of Prediction Performance of Causal Models

Arxiv

0+阅读 · 2021年10月26日

Random matrices based schemes for stable and robust nonparametric and functional regression estimators

Arxiv

0+阅读 · 2021年10月26日

Optimal Bayesian Estimation of a Regression Curve, a Conditional Density and a Conditional Distribution

Arxiv

0+阅读 · 2021年10月26日

The SKIM-FA Kernel: High-Dimensional Variable Selection and Nonlinear Interaction Discovery in Linear Time

Arxiv

0+阅读 · 2021年10月26日

Towards Practical Mean Bounds for Small Samples

Arxiv

0+阅读 · 2021年10月25日

Sufficient reductions in regression with mixed predictors

Arxiv

0+阅读 · 2021年10月25日

Applying Regression Conformal Prediction with Nearest Neighbors to time series data

Arxiv

0+阅读 · 2021年10月25日

Propensity score regression for causal inference with treatment heterogeneity

Arxiv

0+阅读 · 2021年10月25日

Approximate Core for Committee Selection via Multilinear Extension and Market Clearing

Arxiv

0+阅读 · 2021年10月24日

VIP会员

文章信息

相关主题

相关VIP内容

【经典书】线性代数，436页pdf

专知会员服务

78+阅读 · 2021年3月16日

【经典书】线性代数元素，197页pdf

【经典书】线性代数元素，197页pdf

专知会员服务

57+阅读 · 2021年3月4日

最新《图理论》笔记书，98页pdf

最新《图理论》笔记书，98页pdf

专知会员服务

76+阅读 · 2020年12月27日

【干货书】机器学习速查手册，135页pdf

【干货书】机器学习速查手册，135页pdf

专知会员服务

127+阅读 · 2020年11月20日

【Manning新书】现代Java实战，592页pdf

【Manning新书】现代Java实战，592页pdf

专知会员服务

101+阅读 · 2020年5月22日

【ICLR2020】理解非自回归机器翻译中的知识蒸馏（Understanding Knowledge Distillation in Non-autoregressive Machine Translation）

【ICLR2020】理解非自回归机器翻译中的知识蒸馏（Understanding Knowledge Distillation in Non-autoregressive Machine Translation）

专知会员服务

11+阅读 · 2019年12月28日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【新书】Python编程基础，669页pdf

【新书】Python编程基础，669页pdf

专知会员服务

197+阅读 · 2019年10月10日

热门VIP内容

开通专知VIP会员享更多权益服务

【博士论文】在低维和高维空间中分析、建模和转换潜在表征

从无人机到数据：揭示边缘计算作为新作战域

可解释人工智能的基础

大规模视觉模型中的基于提示的适应：综述

相关资讯

【论文笔记】通俗理解少样本文本分类 (Few-Shot Text Classification) (1)

【论文笔记】通俗理解少样本文本分类 (Few-Shot Text Classification) (1)

深度学习自然语言处理

7+阅读 · 2020年4月8日

分布式并行架构Ray介绍

分布式并行架构Ray介绍

CreateAMind

10+阅读 · 2019年8月9日

已删除

将门创投

5+阅读 · 2019年6月28日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

分布式TensorFlow入门指南

分布式TensorFlow入门指南

机器学习研究会

4+阅读 · 2017年11月28日

【论文】变分推断（Variational inference)的总结

【论文】变分推断（Variational inference)的总结

机器学习研究会

39+阅读 · 2017年11月16日

【学习】(Python)SVM数据分类

【学习】(Python)SVM数据分类

机器学习研究会

6+阅读 · 2017年10月15日

【推荐】决策树/随机森林深入解析

【推荐】决策树/随机森林深入解析

机器学习研究会

5+阅读 · 2017年9月21日

相关论文

Inference in Regression Discontinuity Designs with High-Dimensional Covariates

Arxiv

0+阅读 · 2021年10月26日

Sample Selection Bias in Evaluation of Prediction Performance of Causal Models

Arxiv

0+阅读 · 2021年10月26日

Random matrices based schemes for stable and robust nonparametric and functional regression estimators

Arxiv

0+阅读 · 2021年10月26日

Optimal Bayesian Estimation of a Regression Curve, a Conditional Density and a Conditional Distribution

Arxiv

0+阅读 · 2021年10月26日

The SKIM-FA Kernel: High-Dimensional Variable Selection and Nonlinear Interaction Discovery in Linear Time

Arxiv

0+阅读 · 2021年10月26日

Towards Practical Mean Bounds for Small Samples

Arxiv

0+阅读 · 2021年10月25日

Sufficient reductions in regression with mixed predictors

Arxiv

0+阅读 · 2021年10月25日

Applying Regression Conformal Prediction with Nearest Neighbors to time series data

Arxiv

0+阅读 · 2021年10月25日

Propensity score regression for causal inference with treatment heterogeneity

Arxiv

0+阅读 · 2021年10月25日

Approximate Core for Committee Selection via Multilinear Extension and Market Clearing

Arxiv

0+阅读 · 2021年10月24日

微信扫码咨询专知VIP会员