基于小波表示的视觉Transformer组合性探究 (Exploring Compositionality in Vision Transformers using Wavelet Representations) - 专知论文

会员服务 ·

0

表示 · 组合性 · 视觉Transformer · Transformer · 基元 ·

2025 年 12 月 30 日

Exploring Compositionality in Vision Transformers using Wavelet Representations

翻译：基于小波表示的视觉Transformer组合性探究

Akshad Shyam Purushottamdas,Pranav K Nayak,Divya Mehul Rajparia,Deekshith Patel,Yashmitha Gogineni,Konda Reddy Mopuri,Sumohana S. Channappayya

from arxiv, 9 pages, 6 figures

While insights into the workings of the transformer model have largely emerged by analysing their behaviour on language tasks, this work investigates the representations learnt by the Vision Transformer (ViT) encoder through the lens of compositionality. We introduce a framework, analogous to prior work on measuring compositionality in representation learning, to test for compositionality in the ViT encoder. Crucial to drawing this analogy is the Discrete Wavelet Transform (DWT), which is a simple yet effective tool for obtaining input-dependent primitives in the vision setting. By examining the ability of composed representations to reproduce original image representations, we empirically test the extent to which compositionality is respected in the representation space. Our findings show that primitives from a one-level DWT decomposition produce encoder representations that approximately compose in latent space, offering a new perspective on how ViTs structure information.

翻译：尽管对Transformer模型工作机制的洞察主要源于对其在语言任务中行为的分析，本研究通过组合性的视角探究了视觉Transformer（ViT）编码器所学习到的表示。我们引入了一个与先前衡量表示学习中组合性的工作相类似的框架，用以检验ViT编码器中的组合性。建立这种类比的关键在于离散小波变换（DWT），它是一种在视觉场景中获取输入相关基元的简单而有效的工具。通过检验组合表示重构原始图像表示的能力，我们实证测试了表示空间在多大程度上遵循组合性。我们的研究结果表明，来自单层DWT分解的基元所产生的编码器表示在潜在空间中近似可组合，这为理解ViT如何组织信息提供了新的视角。

0

相关内容

【CVPR2024】自然监督下的三维视觉定位与语言规范化的概念学习

【CVPR2024】自然监督下的三维视觉定位与语言规范化的概念学习

专知会员服务

16+阅读 · 2024年5月1日

UTC: 用于视觉对话的任务间对比学习的统一Transformer

UTC: 用于视觉对话的任务间对比学习的统一Transformer

专知会员服务

14+阅读 · 2022年5月4日

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

专知会员服务

68+阅读 · 2022年3月20日

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

专知会员服务

12+阅读 · 2022年3月19日

【伯克利JD Co-Reyes博士论文】建立强化学习算法泛化:从潜在动力学模型到元学习，Building Reinforcement Learning Algorithms that Generalize: From Latent Dynamics Models to Meta-Learning

【伯克利JD Co-Reyes博士论文】建立强化学习算法泛化:从潜在动力学模型到元学习，Building Reinforcement Learning Algorithms that Generalize: From Latent Dynamics Models to Meta-Learning

专知会员服务

45+阅读 · 2022年3月6日

【CVPR 2020 Oral】小样本类增量学习

【CVPR 2020 Oral】小样本类增量学习

专知

20+阅读 · 2020年6月26日

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

专知

13+阅读 · 2020年4月1日

图机器学习 2.2-2.4 Properties of Networks, Random Graph

图机器学习 2.2-2.4 Properties of Networks, Random Graph

图与推荐

10+阅读 · 2020年3月28日

论文浅尝 | Interaction Embeddings for Prediction and Explanation

论文浅尝 | Interaction Embeddings for Prediction and Explanation

开放知识图谱

11+阅读 · 2019年2月1日

CosFace: Large Margin Cosine Loss for Deep Face Recognition论文笔记

CosFace: Large Margin Cosine Loss for Deep Face Recognition论文笔记

统计学习与视觉计算组

44+阅读 · 2018年4月25日

基于各向异性点光源的近场光度学三维重建问题研究

国家自然科学基金

2+阅读 · 2017年12月31日

2D/3D视觉信息融合仿生SLAM关键问题研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于自主学习的Ad hoc Agent序贯决策研究

国家自然科学基金

46+阅读 · 2015年12月31日

面向学术资源的TSD与TDC测度及分析研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于决策模型和预备电位的运动想象BCI研究

国家自然科学基金

3+阅读 · 2015年12月31日

Towards Integrating Uncertainty for Domain-Agnostic Segmentation

Arxiv

0+阅读 · 2025年12月29日

Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition

Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition

Arxiv

0+阅读 · 2025年12月29日

Active Constraint Learning in High Dimensions from Demonstrations

Arxiv

0+阅读 · 2025年12月28日

Poisson-Process Topic Model for Integrating Knowledge from Pre-trained Language Models

Arxiv

0+阅读 · 2025年12月26日

Syntax Is Not Enough: An Empirical Study of Small Transformer Models for Neural Code Repair

Arxiv

0+阅读 · 2025年12月22日

VIP会员

文章信息

相关主题

视觉Transformer

相关VIP内容

【CVPR2024】自然监督下的三维视觉定位与语言规范化的概念学习

【CVPR2024】自然监督下的三维视觉定位与语言规范化的概念学习

专知会员服务

16+阅读 · 2024年5月1日

UTC: 用于视觉对话的任务间对比学习的统一Transformer

UTC: 用于视觉对话的任务间对比学习的统一Transformer

专知会员服务

14+阅读 · 2022年5月4日

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

专知会员服务

68+阅读 · 2022年3月20日

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

【CVPR 2022】基于视觉-语言验证和迭代推理的视觉定位,Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation

专知会员服务

12+阅读 · 2022年3月19日

【伯克利JD Co-Reyes博士论文】建立强化学习算法泛化:从潜在动力学模型到元学习，Building Reinforcement Learning Algorithms that Generalize: From Latent Dynamics Models to Meta-Learning

【伯克利JD Co-Reyes博士论文】建立强化学习算法泛化:从潜在动力学模型到元学习，Building Reinforcement Learning Algorithms that Generalize: From Latent Dynamics Models to Meta-Learning

专知会员服务

45+阅读 · 2022年3月6日

热门VIP内容

开通专知VIP会员享更多权益服务

《运用增强现实技术进行军事任务规划》130页

《高压决策环境中的人机协作》200页博士论文

《2025财年美陆军转型倡议（ATI）部队结构与组织提案》

《探索用于低层级任务区分与分类的转址旁路缓冲》

相关资讯

【CVPR 2020 Oral】小样本类增量学习

【CVPR 2020 Oral】小样本类增量学习

专知

20+阅读 · 2020年6月26日

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

【CVPR2020-旷视】DPGN：分布传播图网络的小样本学习

专知

13+阅读 · 2020年4月1日

图机器学习 2.2-2.4 Properties of Networks, Random Graph

图机器学习 2.2-2.4 Properties of Networks, Random Graph

图与推荐

10+阅读 · 2020年3月28日

论文浅尝 | Interaction Embeddings for Prediction and Explanation

论文浅尝 | Interaction Embeddings for Prediction and Explanation

开放知识图谱

11+阅读 · 2019年2月1日

CosFace: Large Margin Cosine Loss for Deep Face Recognition论文笔记

CosFace: Large Margin Cosine Loss for Deep Face Recognition论文笔记

统计学习与视觉计算组

44+阅读 · 2018年4月25日

相关论文

Towards Integrating Uncertainty for Domain-Agnostic Segmentation

Arxiv

0+阅读 · 2025年12月29日

Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition

Multi-Track Multimodal Learning on iMiGUE: Micro-Gesture and Emotion Recognition

Arxiv

0+阅读 · 2025年12月29日

Active Constraint Learning in High Dimensions from Demonstrations

Arxiv

0+阅读 · 2025年12月28日

Poisson-Process Topic Model for Integrating Knowledge from Pre-trained Language Models

Arxiv

0+阅读 · 2025年12月26日

Syntax Is Not Enough: An Empirical Study of Small Transformer Models for Neural Code Repair

Arxiv

0+阅读 · 2025年12月22日

相关基金

基于各向异性点光源的近场光度学三维重建问题研究

国家自然科学基金

2+阅读 · 2017年12月31日

2D/3D视觉信息融合仿生SLAM关键问题研究

国家自然科学基金

3+阅读 · 2015年12月31日

基于自主学习的Ad hoc Agent序贯决策研究

国家自然科学基金

46+阅读 · 2015年12月31日

面向学术资源的TSD与TDC测度及分析研究

国家自然科学基金

1+阅读 · 2015年12月31日

基于决策模型和预备电位的运动想象BCI研究

国家自然科学基金

3+阅读 · 2015年12月31日

微信扫码咨询专知VIP会员