探索容忍与非容忍分发测试之间的差距 (Exploring the Gap between Tolerant and Non-tolerant Distribution Testing)

The framework of distribution testing is currently ubiquitous in the field of property testing. In this model, the input is a probability distribution accessible via independently drawn samples from an oracle. The testing task is to distinguish a distribution that satisfies some property from a distribution that is far from satisfying it in the $\ell_1$ distance. The task of tolerant testing imposes a further restriction, that distributions close to satisfying the property are also accepted. This work focuses on the connection of the sample complexities of non-tolerant ("traditional") testing of distributions and tolerant testing thereof. When limiting our scope to label-invariant (symmetric) properties of distribution, we prove that the gap is at most quadratic. Conversely, the property of being the uniform distribution is indeed known to have an almost-quadratic gap. When moving to general, not necessarily label-invariant properties, the situation is more complicated, and we show some partial results. We show that if a property requires the distributions to be non-concentrated, then it cannot be non-tolerantly tested with $o(\sqrt{n})$ many samples, where $n$ denotes the universe size. Clearly, this implies at most a quadratic gap, because a distribution can be learned (and hence tolerantly tested against any property) using $\mathcal{O}(n)$ many samples. Being non-concentrated is a strong requirement on the property, as we also prove a close to linear lower bound against their tolerant tests. To provide evidence for other general cases (where the properties are not necessarily label-invariant), we show that if an input distribution is very concentrated, in the sense that it is mostly supported on a subset of size $s$ of the universe, then it can be learned using only $\mathcal{O}(s)$ many samples. The learning procedure adapts to the input, and works without knowing $s$ in advance.

翻译：分布测试框架目前无处不在属性测试领域 { 。在这个模型中, 输入是一个可以通过独立提取的样本获取的概率分布。测试的任务是将满足某些属性的分布与远不能满足该属性的分布区别开来 $\ ell_ 1美元距离。宽容测试的任务进一步施加了限制, 分配接近于满足属性的分布也被接受。这项工作侧重于不宽容( “ 传统” ) 的分布测试及其宽容测试的样本复杂性的连接。当将我们的范围限制在分配的标签- 变量( 符号) 属性属性的特性时, 我们只能通过独立提取的样本( 符号) 。相反, 统一分布的属性确实存在几乎不满足该属性的距离。当移动到普通时, 情况会更加复杂。我们显示, 如果某个属性需要使用不宽容的分布, 则无法用美元( ) ( 符号值) 的( ) 直径( ) 直径( ) 值) 来进行不宽容的测试, 直径( 值) 值) 和最接近的样本的颜色的特性的特性的大小( 。显示美元。