用于对象发现发现的复杂值的自动编码器Name (Complex-Valued Autoencoders for Object Discovery)

Object-centric representations form the basis of human perception and enable us to reason about the world and to systematically generalize to new settings. Currently, most machine learning work on unsupervised object discovery focuses on slot-based approaches, which explicitly separate the latent representations of individual objects. While the result is easily interpretable, it usually requires the design of involved architectures. In contrast to this, we propose a distributed approach to object-centric representations: the Complex AutoEncoder. Following a coding scheme theorized to underlie object representations in biological neurons, its complex-valued activations represent two messages: their magnitudes express the presence of a feature, while the relative phase differences between neurons express which features should be bound together to create joint object representations. We show that this simple and efficient approach achieves better reconstruction performance than an equivalent real-valued autoencoder on simple multi-object datasets. Additionally, we show that it achieves competitive unsupervised object discovery performance to a SlotAttention model on two datasets, and manages to disentangle objects in a third dataset where SlotAttention fails - all while being 7-70 times faster to train.

翻译：以物体为中心的表达方式构成了人类感知的基础, 并使我们能够了解世界, 并系统化地概括到新的设置。目前, 大多数关于不受监督的物体发现机器学习工作都集中在基于时间档的方法上, 这种方法明确区分了单个物体的潜在表现。虽然结果很容易解释, 但通常需要设计相关的结构。与此相反, 我们建议对以物体为中心的表达方式采取分布式方法: 复杂自动编码器。在对生物神经体中物体表示表示的物体表示结构进行编码后, 其复杂价值的激活代表了两种信息: 它们的数量表示一个特性的存在, 而神经体之间的相对阶段差异表示哪些特性应该捆绑在一起以创建共同的物体表示。我们表明, 这种简单而有效的方法比简单多对象数据集上一个等效的、实际价值的自动编码器的重建性效果要好。此外, 我们显示, 在两个数据集中, 它能取得竞争性的、不超强的物体发现性物体发现性能, 也就是两个SlotAnyaction 模型, 并且能够将第三个数据设置中的物体分解, 即为7-70 快速的训练。

相关内容

自编码器

关注 140

自动编码器是一种人工神经网络，用于以无监督的方式学习有效的数据编码。自动编码器的目的是通过训练网络忽略信号“噪声”来学习一组数据的表示（编码），通常用于降维。与简化方面一起，学习了重构方面，在此，自动编码器尝试从简化编码中生成尽可能接近其原始输入的表示形式，从而得到其名称。基本模型存在几种变体，其目的是迫使学习的输入表示形式具有有用的属性。自动编码器可有效地解决许多应用问题，从面部识别到获取单词的语义。

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日