神经架构表征学习中的综合属性预测：NAR-Former (NAR-Former: Neural Architecture Representation Learning towards Holistic Attributes Prediction)

With the wide and deep adoption of deep learning models in real applications, there is an increasing need to model and learn the representations of the neural networks themselves. These models can be used to estimate attributes of different neural network architectures such as the accuracy and latency, without running the actual training or inference tasks. In this paper, we propose a neural architecture representation model that can be used to estimate these attributes holistically. Specifically, we first propose a simple and effective tokenizer to encode both the operation and topology information of a neural network into a single sequence. Then, we design a multi-stage fusion transformer to build a compact vector representation from the converted sequence. For efficient model training, we further propose an information flow consistency augmentation and correspondingly design an architecture consistency loss, which brings more benefits with less augmentation samples compared with previous random augmentation strategies. Experiment results on NAS-Bench-101, NAS-Bench-201, DARTS search space and NNLQP show that our proposed framework can be used to predict the aforementioned latency and accuracy attributes of both cell architectures and whole deep neural networks, and achieves promising performance. Code is available at https://github.com/yuny220/NAR-Former.

翻译：随着深度学习模型在实际应用中的广泛使用，人们越来越需要对神经网络本身进行建模和学习表征模型。这些模型可以用来估计不同神经网络架构的属性，例如准确率和延迟，而无需运行实际的训练或推理任务。在本文中，我们提出了一种神经架构表征模型，用于综合预测这些属性。具体而言，我们首先提出了一种简单而有效的标记器，将神经网络的运算和拓扑信息编码成一个单独的序列。然后，我们设计了一个多阶段融合变压器，从转换后的序列构建一个紧凑的向量表示。为了实现高效的模型训练，我们进一步提出了信息流程一致性增强，相应地设计了一种架构一致性损失，与以前的随机增强策略相比，它在使用更少增强样本时带来了更多的好处。在NAS-Bench-101、NAS-Bench-201、DARTS搜索空间和NNLQP上的实验结果表明，我们提出的框架可以用于预测单个细胞和整个深度神经网络的延迟和准确率属性，并实现了良好的性能。代码可在https://github.com/yuny220/NAR-Former上找到。

相关内容

Neural Networks

关注 1648

神经网络（Neural Networks）是世界上三个最古老的神经建模学会的档案期刊:国际神经网络学会(INNS)、欧洲神经网络学会(ENNS)和日本神经网络学会(JNNS)。神经网络提供了一个论坛，以发展和培育一个国际社会的学者和实践者感兴趣的所有方面的神经网络和相关方法的计算智能。神经网络欢迎高质量论文的提交，有助于全面的神经网络研究，从行为和大脑建模，学习算法，通过数学和计算分析，系统的工程和技术应用，大量使用神经网络的概念和技术。这一独特而广泛的范围促进了生物和技术研究之间的思想交流，并有助于促进对生物启发的计算智能感兴趣的跨学科社区的发展。因此，神经网络编委会代表的专家领域包括心理学，神经生物学，计算机科学，工程，数学，物理。该杂志发表文章、信件和评论以及给编辑的信件、社论、时事、软件调查和专利信息。文章发表在五个部分之一:认知科学，神经科学，学习系统，数学和计算分析、工程和应用。官网地址：http://dblp.uni-trier.de/db/journals/nn/

【CMU博士论文】神经架构搜索的搜索算法和搜索空间，141页pdf

专知会员服务

38+阅读 · 2022年12月7日

【IJCAI2021】User-as-Graph: 基于异构图池化的新闻推荐用户建模

专知会员服务

23+阅读 · 2021年8月25日

【ICML2020-斯坦福Facebook-何恺明】神经网络图结构，Graph Structure of Neural Networks

专知会员服务

57+阅读 · 2020年7月14日

【ICML2020】深度神经网络置信感知学习，Conﬁdence-Aware Learning for Deep Neural Networks

专知会员服务

74+阅读 · 2020年7月6日