翻译后的标题: (WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation) - 专知论文

会员服务 ·

0

弱监督语义分割 · 监督 · 视觉Transformer · 语义分割 · 分割 ·

2023 年 4 月 3 日

WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation

翻译：翻译后的标题:

Lianghui Zhu,Yingyue Li,Jieming Fang,Yan Liu,Hao Xin,Wenyu Liu,Xinggang Wang

from arxiv, 20 pages, 11 figures

This paper explores the properties of the plain Vision Transformer (ViT) for Weakly-supervised Semantic Segmentation (WSSS). The class activation map (CAM) is of critical importance for understanding a classification network and launching WSSS. We observe that different attention heads of ViT focus on different image areas. Thus a novel weight-based method is proposed to end-to-end estimate the importance of attention heads, while the self-attention maps are adaptively fused for high-quality CAM results that tend to have more complete objects. Besides, we propose a ViT-based gradient clipping decoder for online retraining with the CAM results to complete the WSSS task. We name this plain Transformer-based Weakly-supervised learning framework WeakTr. It achieves the state-of-the-art WSSS performance on standard benchmarks, i.e., 78.4% mIoU on the val set of PASCAL VOC 2012 and 50.3% mIoU on the val set of COCO 2014. Code is available at https://github.com/hustvl/WeakTr.

翻译：WeakTr：探索平凡的视觉Transformer在弱监督语义分割中的应用翻译后的摘要：本文探讨了平凡的视觉Transformer（ViT）在弱监督语义分割（WSSS）中的应用。类激活图(CAM)对于理解分类网络和启动WSSS至关重要。我们观察到ViT的不同注意力头集中于不同的图像区域。因此，提出了一种基于权重的方法，以端到端地估计注意力头的重要性，同时自适应地融合自注意力图以实现具有更完整对象的高质量CAM结果。此外，我们提出了一种基于ViT的梯度裁剪解码器，通过CAM结果进行在线再训练以完成WSSS任务。我们将基于平凡Transformer的弱监督学习框架称为WeakTr。它在标准基准测试中实现了最先进的WSSS性能，即在PASCAL VOC 2012的val集上为78.4％ mIoU，在COCO 2014的val集上为50.3％ mIoU。代码可在https://github.com/hustvl/WeakTr上获得。

0

相关内容

弱监督语义分割

弱监督语义分割

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

专知会员服务

24+阅读 · 2023年4月9日

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

专知会员服务

21+阅读 · 2022年3月18日

【CVPR2022】弱监督语义分割的类重新激活图

【CVPR2022】弱监督语义分割的类重新激活图

专知会员服务

17+阅读 · 2022年3月7日

【CVPR2021】基于Transformers 从序列到序列的角度重新思考语义分割

【CVPR2021】基于Transformers 从序列到序列的角度重新思考语义分割

专知会员服务

44+阅读 · 2021年3月15日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

专知会员服务

92+阅读 · 2019年12月22日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

【泡泡汇总】CVPR2019 SLAM Paperlist

【泡泡汇总】CVPR2019 SLAM Paperlist

泡泡机器人SLAM

14+阅读 · 2019年6月12日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

泡泡机器人SLAM

22+阅读 · 2018年12月4日

【泡泡机器人】ECCV2018之SLAM最新前沿动态（附文章链接和代码链接）

【泡泡机器人】ECCV2018之SLAM最新前沿动态（附文章链接和代码链接）

泡泡机器人SLAM

38+阅读 · 2018年9月23日

《pyramid Attention Network for Semantic Segmentation》

《pyramid Attention Network for Semantic Segmentation》

统计学习与视觉计算组

44+阅读 · 2018年8月30日

【泡泡一分钟】端到端的弱监督语义对齐

【泡泡一分钟】端到端的弱监督语义对齐

泡泡机器人SLAM

53+阅读 · 2018年4月5日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

【论文推荐】最新5篇目标检测相关论文——显著目标检测、弱监督One-Shot检测、多框检测器、携带物体检测、假彩色图像检测

【论文推荐】最新5篇目标检测相关论文——显著目标检测、弱监督One-Shot检测、多框检测器、携带物体检测、假彩色图像检测

专知

74+阅读 · 2018年1月16日

双靶点干预对糖尿病视网膜病变中“血管-纤维化开关”的调控

国家自然科学基金

1+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

多组分量子点共敏化太阳能电池的制备与光电性能

国家自然科学基金

0+阅读 · 2013年12月31日

前体mRNA剪切因子SF1磷酸化的结构与功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

紫外纳米线/可见光量子点型复合光催化材料的构筑及CO2转化性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

PPARγ-1SUMO化修饰在高（血）糖诱导血管内皮胰岛素抵抗中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

玉米幼苗干旱胁迫应答NAC转录因子基因的筛选和鉴定

国家自然科学基金

0+阅读 · 2012年12月31日

胡杨叶形可塑性与生理生态适应机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

TRPC6在VEGF调节新生血管形成中的作用及机制

国家自然科学基金

0+阅读 · 2008年12月31日

丛枝菌根真菌与宿主植物相互作用的相关基因研究

国家自然科学基金

0+阅读 · 2008年12月31日

Prototype Adaption and Projection for Few- and Zero-shot 3D Point Cloud Semantic Segmentation

Arxiv

0+阅读 · 2023年5月23日

MOTRv3: Release-Fetch Supervision for End-to-End Multi-Object Tracking

Arxiv

0+阅读 · 2023年5月23日

Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering

Arxiv

0+阅读 · 2023年5月23日

Object Segmentation by Mining Cross-Modal Semantics

Arxiv

0+阅读 · 2023年5月23日

Exploring Train and Test-Time Augmentations for Audio-Language Learning

Arxiv

0+阅读 · 2023年5月23日

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

Arxiv

0+阅读 · 2023年5月22日

Uncertainty-based Detection of Adversarial Attacks in Semantic Segmentation

Arxiv

0+阅读 · 2023年5月22日

Spatiotemporal Attention-based Semantic Compression for Real-time Video Recognition

Arxiv

0+阅读 · 2023年5月22日

Enhancing Transformer Backbone for Egocentric Video Action Segmentation

Arxiv

0+阅读 · 2023年5月19日

Temporal Relational Modeling with Self-Supervision for Action Segmentation

Arxiv

13+阅读 · 2020年12月14日

VIP会员

文章信息

相关主题

弱监督语义分割

视觉Transformer

相关VIP内容

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

【CVPR2023】Mask3D:通过学习掩码3D先验对2D视觉transformer进行预训练

专知会员服务

24+阅读 · 2023年4月9日

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

【Hugging Face】使用自定义数据集微调语义分割模型，Fine-Tune a Semantic Segmentation Model with a Custom Dataset

专知会员服务

21+阅读 · 2022年3月18日

【CVPR2022】弱监督语义分割的类重新激活图

【CVPR2022】弱监督语义分割的类重新激活图

专知会员服务

17+阅读 · 2022年3月7日

【CVPR2021】基于Transformers 从序列到序列的角度重新思考语义分割

【CVPR2021】基于Transformers 从序列到序列的角度重新思考语义分割

专知会员服务

44+阅读 · 2021年3月15日

【CVPR2021】用Transformers无监督预训练进行目标检测

【CVPR2021】用Transformers无监督预训练进行目标检测

专知会员服务

58+阅读 · 2021年3月3日

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

【CVPR2020-中科院计算所】弱监督语义分割的自监督等价注意力机制，Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

专知会员服务

76+阅读 · 2020年4月10日

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

抢鲜看！13篇CVPR2020论文链接/开源代码/解读

专知会员服务

50+阅读 · 2020年2月26日

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

【Google ICLR2020论文】嵌入式大规模检索的预训练任务，Pre-training Tasks for Embedding-based Large-scale Retrieval

专知会员服务

28+阅读 · 2020年2月12日

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

【微软研究院】IMAGEBERT: CROSS-MODAL PRE-TRAINING WITH LARGE-SCALE WEAK-SUPERVISED IMAGE-TEXT DATA

专知会员服务

43+阅读 · 2020年1月28日

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

【AAAI2020】多模态注意力语义图嵌入多标签分类（Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification）

专知会员服务

92+阅读 · 2019年12月22日

热门VIP内容

开通专知VIP会员享更多权益服务

【牛津博士论文】零样本强化学习综述

《美军条令：陆军指挥官与规划人员地理空间指南》60页

战术边缘指挥控制：防务面临的核心挑战

迈向开放世界检测：综述

相关资讯

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

【泡泡汇总】CVPR2019 SLAM Paperlist

【泡泡汇总】CVPR2019 SLAM Paperlist

泡泡机器人SLAM

14+阅读 · 2019年6月12日

BERT/Transformer/迁移学习NLP资源大列表

BERT/Transformer/迁移学习NLP资源大列表

专知

19+阅读 · 2019年6月9日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

【泡泡一分钟】用于RGBD语义分割的三维图神经网络(ICCV2017-546)

泡泡机器人SLAM

22+阅读 · 2018年12月4日

【泡泡机器人】ECCV2018之SLAM最新前沿动态（附文章链接和代码链接）

【泡泡机器人】ECCV2018之SLAM最新前沿动态（附文章链接和代码链接）

泡泡机器人SLAM

38+阅读 · 2018年9月23日

《pyramid Attention Network for Semantic Segmentation》

《pyramid Attention Network for Semantic Segmentation》

统计学习与视觉计算组

44+阅读 · 2018年8月30日

【泡泡一分钟】端到端的弱监督语义对齐

【泡泡一分钟】端到端的弱监督语义对齐

泡泡机器人SLAM

53+阅读 · 2018年4月5日

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

【论文推荐】最新5篇图像分割（Image Segmentation）相关论文—多重假设、超像素分割、自监督、图、生成对抗网络

专知

27+阅读 · 2018年2月7日

【论文推荐】最新5篇目标检测相关论文——显著目标检测、弱监督One-Shot检测、多框检测器、携带物体检测、假彩色图像检测

【论文推荐】最新5篇目标检测相关论文——显著目标检测、弱监督One-Shot检测、多框检测器、携带物体检测、假彩色图像检测

专知

74+阅读 · 2018年1月16日

相关论文

Prototype Adaption and Projection for Few- and Zero-shot 3D Point Cloud Semantic Segmentation

Arxiv

0+阅读 · 2023年5月23日

MOTRv3: Release-Fetch Supervision for End-to-End Multi-Object Tracking

Arxiv

0+阅读 · 2023年5月23日

Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering

Arxiv

0+阅读 · 2023年5月23日

Object Segmentation by Mining Cross-Modal Semantics

Arxiv

0+阅读 · 2023年5月23日

Exploring Train and Test-Time Augmentations for Audio-Language Learning

Arxiv

0+阅读 · 2023年5月23日

HGFormer: Hierarchical Grouping Transformer for Domain Generalized Semantic Segmentation

Arxiv

0+阅读 · 2023年5月22日

Uncertainty-based Detection of Adversarial Attacks in Semantic Segmentation

Arxiv

0+阅读 · 2023年5月22日

Spatiotemporal Attention-based Semantic Compression for Real-time Video Recognition

Arxiv

0+阅读 · 2023年5月22日

Enhancing Transformer Backbone for Egocentric Video Action Segmentation

Arxiv

0+阅读 · 2023年5月19日

Temporal Relational Modeling with Self-Supervision for Action Segmentation

Arxiv

13+阅读 · 2020年12月14日

相关基金

双靶点干预对糖尿病视网膜病变中“血管-纤维化开关”的调控

国家自然科学基金

1+阅读 · 2014年12月31日

PPAR β/δ基因在结直肠癌血管生成调控中的作用及分子机理

国家自然科学基金

2+阅读 · 2014年12月31日

多组分量子点共敏化太阳能电池的制备与光电性能

国家自然科学基金

0+阅读 · 2013年12月31日

前体mRNA剪切因子SF1磷酸化的结构与功能研究

国家自然科学基金

0+阅读 · 2013年12月31日

紫外纳米线/可见光量子点型复合光催化材料的构筑及CO2转化性能研究

国家自然科学基金

0+阅读 · 2012年12月31日

PPARγ-1SUMO化修饰在高（血）糖诱导血管内皮胰岛素抵抗中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

玉米幼苗干旱胁迫应答NAC转录因子基因的筛选和鉴定

国家自然科学基金

0+阅读 · 2012年12月31日

胡杨叶形可塑性与生理生态适应机理研究

国家自然科学基金

0+阅读 · 2009年12月31日

TRPC6在VEGF调节新生血管形成中的作用及机制

国家自然科学基金

0+阅读 · 2008年12月31日

丛枝菌根真菌与宿主植物相互作用的相关基因研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员