多模型估计的自助法以考虑模型选择 (Bootstrapping multiple systems estimates to account for model selection) - 专知论文

会员服务 ·

0

模型选择 · 模型估计 · 多模型 · 准则 · 多模 ·

2023 年 3 月 31 日

Bootstrapping multiple systems estimates to account for model selection

翻译：多模型估计的自助法以考虑模型选择

Bernard W. Silverman,Kyle Vincent,Lax Chan

from arxiv, 21 pages, 5 figures, 6 tables

Multiple systems estimation is a standard approach to quantifying hidden populations where data sources are based on lists of known cases. A typical modelling approach is to fit a Poisson loglinear model to the numbers of cases observed in each possible combination of the lists. It is necessary to decide which interaction parameters to include in the model, and information criterion approaches are often used for model selection. Difficulties in the context of multiple systems estimation may arise due to sparse or nil counts based on the intersection of lists, and care must be taken when information criterion approaches are used for model selection due to issues relating to the existence of estimates and identifiability of the model. Confidence intervals are often reported conditional on the model selected, providing an over-optimistic impression of the accuracy of the estimation. A bootstrap approach is a natural way to account for the model selection procedure. However, because the model selection step has to be carried out for every bootstrap replication, there may be a high or even prohibitive computational burden. We explore the merit of modifying the model selection procedure in the bootstrap to look only among a subset of models, chosen on the basis of their information criterion score on the original data. This provides large computational gains with little apparent effect on inference. Another model selection approach considered and investigated is a downhill search approach among models, possibly with multiple starting points.

翻译：多模型估计是一种量化基于已知案例列表的隐藏人口的标准方法。典型的建模方法是对每个列表组合中观察到的案例数量拟合泊松对数线性模型。必须决定在模型中包括哪些交互参数，并且通常使用信息准则方法进行模型选择。由于基于列表交集的稀疏或零计数而导致的困难可能会出现在多模型估计的情况下，并且由于存在估计和模型可识别性问题，因此在使用信息准则方法进行模型选择时必须小心。通常在选定模型的条件下报告置信区间，从而提供有关估计准确性的过分乐观印象。引导法是解决模型选择过程的自然方法。但是，由于必须为每个引导式重复进行模型选择步骤，因此可能存在高甚至不能承受的计算负担。我们探讨了修改引导式中的模型选择过程的优点，以仅在基于原始数据的信息准则得分选择的一组模型中查找。这提供了大的计算收益，对推理几乎没有影响。还考虑并研究了一种下山式搜索方法，以在模型之间进行选择，可能具有多个起点。

0

相关内容

模型选择

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

【简明书】数学，统计和机器学习的动手入门，57页pdf，A Hands-On Introduction to Math, Stats, and Machine Learning

【简明书】数学，统计和机器学习的动手入门，57页pdf，A Hands-On Introduction to Math, Stats, and Machine Learning

专知会员服务

43+阅读 · 2022年2月26日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【论文推荐】文本摘要简述

【论文推荐】文本摘要简述

专知会员服务

69+阅读 · 2020年7月20日

对话推荐系统综述论文，35页pdf，A Survey on Conversational Recommender Systems

对话推荐系统综述论文，35页pdf，A Survey on Conversational Recommender Systems

专知会员服务

117+阅读 · 2020年4月3日

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

专知会员服务

32+阅读 · 2020年2月1日

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

专知会员服务

20+阅读 · 2020年1月7日

【AAAI2020接受论文】预测性参与:开放领域对话系统自动评估的有效指标（Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems）

【AAAI2020接受论文】预测性参与:开放领域对话系统自动评估的有效指标（Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems）

专知会员服务

14+阅读 · 2019年11月15日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

不再让CPU和总线拖后腿：Exafunction让GPU跑的更快！

不再让CPU和总线拖后腿：Exafunction让GPU跑的更快！

机器之心

0+阅读 · 2022年10月7日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

缺失数据统计分析，第三版，462页pdf

缺失数据统计分析，第三版，462页pdf

专知

48+阅读 · 2020年2月28日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【SIGIR2018】五篇对抗训练文章

【SIGIR2018】五篇对抗训练文章

专知

12+阅读 · 2018年7月9日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

方差正则化的分类模型选择方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于复合分位数回归和最大秩相关想法的ROC回归曲线估计

国家自然科学基金

0+阅读 · 2013年12月31日

超高维半参数回归模型的结构识别和变量选择问题研究

国家自然科学基金

0+阅读 · 2013年12月31日

复杂制造过程中轮廓数据监控方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

连续时间马氏决策过程均值-方差优化问题的研究

国家自然科学基金

0+阅读 · 2012年12月31日

含有缺失值的纵向数据回归模型的稳健推断

国家自然科学基金

3+阅读 · 2012年12月31日

半参数回归分析的随机函数法及其高维情形

国家自然科学基金

2+阅读 · 2012年12月31日

基于参数和半参数回归模型的小区域估计问题研究

国家自然科学基金

0+阅读 · 2012年12月31日

用多重假设检验方法来研究方差变点问题

国家自然科学基金

0+阅读 · 2009年12月31日

区间删失数据下竞争风险模型研究

国家自然科学基金

0+阅读 · 2008年12月31日

Lightweight Online Learning for Sets of Related Problems in Automated Reasoning

Arxiv

0+阅读 · 2023年5月22日

Approximating a RUM from Distributions on k-Slates

Arxiv

0+阅读 · 2023年5月22日

Quantifying the effect of X-ray scattering for data generation in real-time defect detection

Arxiv

0+阅读 · 2023年5月22日

A parametric distribution for exact post-selection inference with data carving

Arxiv

0+阅读 · 2023年5月21日

Precise Unbiased Estimation in Randomized Experiments using Auxiliary Observational Data

Arxiv

0+阅读 · 2023年5月19日

Bayesian inference for misspecified generative models

Arxiv

0+阅读 · 2023年5月19日

Towards Intersectional Moderation: An Alternative Model of Moderation Built on Care and Power

Arxiv

0+阅读 · 2023年5月18日

DGPO: Discovering Multiple Strategies with Diversity-Guided Policy Optimization

Arxiv

0+阅读 · 2023年5月18日

Spectral Change Point Estimation for High Dimensional Time Series by Sparse Tensor Decomposition

Arxiv

0+阅读 · 2023年5月18日

Using Perturbation to Improve Goodness-of-Fit Tests based on Kernelized Stein Discrepancy

Arxiv

0+阅读 · 2023年5月17日

VIP会员

文章信息

相关主题

相关VIP内容

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

宾夕法尼亚大学最新《不确定性估计》课程笔记，134页pdf，附Slides

专知会员服务

49+阅读 · 2022年11月13日

【简明书】数学，统计和机器学习的动手入门，57页pdf，A Hands-On Introduction to Math, Stats, and Machine Learning

【简明书】数学，统计和机器学习的动手入门，57页pdf，A Hands-On Introduction to Math, Stats, and Machine Learning

专知会员服务

43+阅读 · 2022年2月26日

INRIA 最新《机器学习理论》课程笔记，176页pdf

专知会员服务

51+阅读 · 2020年12月14日

【论文推荐】文本摘要简述

【论文推荐】文本摘要简述

专知会员服务

69+阅读 · 2020年7月20日

对话推荐系统综述论文，35页pdf，A Survey on Conversational Recommender Systems

对话推荐系统综述论文，35页pdf，A Survey on Conversational Recommender Systems

专知会员服务

117+阅读 · 2020年4月3日

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

【论文推荐WWW2020-UIUC】修正排序系统中的选择偏差：Correcting for Selection Bias in Learning-to-rank Systems

专知会员服务

32+阅读 · 2020年2月1日

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

【独立研究者I-Sheng Yang论文】因果机器学习损失函数（A Loss-Function for Causal Machine-Learning）

专知会员服务

20+阅读 · 2020年1月7日

【AAAI2020接受论文】预测性参与:开放领域对话系统自动评估的有效指标（Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems）

【AAAI2020接受论文】预测性参与:开放领域对话系统自动评估的有效指标（Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems）

专知会员服务

14+阅读 · 2019年11月15日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

热门VIP内容

开通专知VIP会员享更多权益服务

【CMU博士论文】基础模型训练中网络规模数据的负责任与高效使用

《俄乌战争背景下俄罗斯的战略性海军分析（2022-2025年）》最新100页报告

人工智能时代背景下的未来海战

相关资讯

不再让CPU和总线拖后腿：Exafunction让GPU跑的更快！

不再让CPU和总线拖后腿：Exafunction让GPU跑的更快！

机器之心

0+阅读 · 2022年10月7日

VCIP 2022 Call for Demos

VCIP 2022 Call for Demos

CCF多媒体专委会

1+阅读 · 2022年6月6日

缺失数据统计分析，第三版，462页pdf

缺失数据统计分析，第三版，462页pdf

专知

48+阅读 · 2020年2月28日

局部学习的特征选择：Local-Learning-Based Feature Selection

局部学习的特征选择：Local-Learning-Based Feature Selection

我爱读PAMI

14+阅读 · 2019年9月20日

强化学习三篇论文避免遗忘等

强化学习三篇论文避免遗忘等

CreateAMind

20+阅读 · 2019年5月24日

逆强化学习-学习人先验的动机

逆强化学习-学习人先验的动机

CreateAMind

16+阅读 · 2019年1月18日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【SIGIR2018】五篇对抗训练文章

【SIGIR2018】五篇对抗训练文章

专知

12+阅读 · 2018年7月9日

【推荐】SVM实例教程

【推荐】SVM实例教程

机器学习研究会

17+阅读 · 2017年8月26日

强化学习族谱

强化学习族谱

CreateAMind

26+阅读 · 2017年8月2日

相关论文

Lightweight Online Learning for Sets of Related Problems in Automated Reasoning

Arxiv

0+阅读 · 2023年5月22日

Approximating a RUM from Distributions on k-Slates

Arxiv

0+阅读 · 2023年5月22日

Quantifying the effect of X-ray scattering for data generation in real-time defect detection

Arxiv

0+阅读 · 2023年5月22日

A parametric distribution for exact post-selection inference with data carving

Arxiv

0+阅读 · 2023年5月21日

Precise Unbiased Estimation in Randomized Experiments using Auxiliary Observational Data

Arxiv

0+阅读 · 2023年5月19日

Bayesian inference for misspecified generative models

Arxiv

0+阅读 · 2023年5月19日

Towards Intersectional Moderation: An Alternative Model of Moderation Built on Care and Power

Arxiv

0+阅读 · 2023年5月18日

DGPO: Discovering Multiple Strategies with Diversity-Guided Policy Optimization

Arxiv

0+阅读 · 2023年5月18日

Spectral Change Point Estimation for High Dimensional Time Series by Sparse Tensor Decomposition

Arxiv

0+阅读 · 2023年5月18日

Using Perturbation to Improve Goodness-of-Fit Tests based on Kernelized Stein Discrepancy

Arxiv

0+阅读 · 2023年5月17日

相关基金

方差正则化的分类模型选择方法研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于复合分位数回归和最大秩相关想法的ROC回归曲线估计

国家自然科学基金

0+阅读 · 2013年12月31日

超高维半参数回归模型的结构识别和变量选择问题研究

国家自然科学基金

0+阅读 · 2013年12月31日

复杂制造过程中轮廓数据监控方法研究

国家自然科学基金

1+阅读 · 2013年12月31日

连续时间马氏决策过程均值-方差优化问题的研究

国家自然科学基金

0+阅读 · 2012年12月31日

含有缺失值的纵向数据回归模型的稳健推断

国家自然科学基金

3+阅读 · 2012年12月31日

半参数回归分析的随机函数法及其高维情形

国家自然科学基金

2+阅读 · 2012年12月31日

基于参数和半参数回归模型的小区域估计问题研究

国家自然科学基金

0+阅读 · 2012年12月31日

用多重假设检验方法来研究方差变点问题

国家自然科学基金

0+阅读 · 2009年12月31日

区间删失数据下竞争风险模型研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员