从树集成模型中提取可解释模型：计算与统计视角 (Extracting Interpretable Models from Tree Ensembles: Computational and Statistical Perspectives)

Tree ensembles are non-parametric methods widely recognized for their accuracy and ability to capture complex interactions. While these models excel at prediction, they are difficult to interpret and may fail to uncover useful relationships in the data. We propose an estimator to extract compact sets of decision rules from tree ensembles. The extracted models are accurate and can be manually examined to reveal relationships between the predictors and the response. A key novelty of our estimator is the flexibility to jointly control the number of rules extracted and the interaction depth of each rule, which improves accuracy. We develop a tailored exact algorithm to efficiently solve optimization problems underlying our estimator and an approximate algorithm for computing regularization paths, sequences of solutions that correspond to varying model sizes. We also establish novel non-asymptotic prediction error bounds for our proposed approach, comparing it to an oracle that chooses the best data-dependent linear combination of the rules in the ensemble subject to the same complexity constraint as our estimator. The bounds illustrate that the large-sample predictive performance of our estimator is on par with that of the oracle. Through experiments, we demonstrate that our estimator outperforms existing algorithms for rule extraction.

翻译：树集成模型作为非参数方法，因其预测精度高且能捕捉复杂交互作用而广受认可。尽管这些模型在预测方面表现优异，但其可解释性较差，且可能无法揭示数据中有用的关联关系。本文提出一种估计器，用于从树集成模型中提取紧凑的决策规则集。所提取的模型不仅精度高，还可通过人工检查来揭示预测变量与响应变量之间的关系。该估计器的关键创新在于能够灵活地联合控制提取规则的数量及各规则的交互深度，从而提升模型精度。我们开发了定制化的精确算法来高效求解估计器所对应的优化问题，并设计了近似算法用于计算正则化路径——即对应不同模型规模的解序列。此外，我们为所提方法建立了新颖的非渐近预测误差界，将其与在相同复杂度约束下选择集成规则最优数据依赖线性组合的预言机估计器进行比较。误差界表明，在大样本情况下，本估计器的预测性能与预言机估计器相当。通过实验验证，本估计器在规则提取任务上优于现有算法。