GRIT: 产生用于了解物体的区域到文字变换器 (GRiT: A Generative Region-to-text Transformer for Object Understanding) - 专知论文

会员服务 ·

0

可理解性 · 目标检测 · 变换 · HTTPS · SimPLe ·

2022 年 12 月 1 日

GRiT: A Generative Region-to-text Transformer for Object Understanding

翻译：GRIT: 产生用于了解物体的区域到文字变换器

Jialian Wu,Jianfeng Wang,Zhengyuan Yang,Zhe Gan,Zicheng Liu,Junsong Yuan,Lijuan Wang

This paper presents a Generative RegIon-to-Text transformer, GRiT, for object understanding. The spirit of GRiT is to formulate object understanding as <region, text> pairs, where region locates objects and text describes objects. For example, the text in object detection denotes class names while that in dense captioning refers to descriptive sentences. Specifically, GRiT consists of a visual encoder to extract image features, a foreground object extractor to localize objects, and a text decoder to generate open-set object descriptions. With the same model architecture, GRiT can understand objects via not only simple nouns, but also rich descriptive sentences including object attributes or actions. Experimentally, we apply GRiT to object detection and dense captioning tasks. GRiT achieves 60.4 AP on COCO 2017 test-dev for object detection and 15.5 mAP on Visual Genome for dense captioning. Code is available at https://github.com/JialianW/GRiT

翻译：本文展示了用于对象理解的生成 Region- Text 变压器, GRIT 。 GRIT 的精神是将对象理解设定为 < 区域, 文本 > 配对, 区域定位对象和文字描述对象。例如, 物体探测中的文字表示类名, 而密集字幕中则指描述性句。具体地说, GRIT 包含用于提取图像特征的视觉编码器、定位对象的地表物体提取器、生成开立对象描述的文本解码器。在同一模型结构下, GRIT 不仅可以通过简单的名词来理解对象,还可以通过包括对象属性或动作在内的内容丰富的描述性句来理解对象。实验性地说, 我们将GRIT 应用于对象探测和密集字幕任务。 GRIT 在 CO 2017 用于物体探测的测试- dev 上实现了60.4 AP, 在视觉基因组用于密集说明的15.5 mAP 。代码见 https://github.com/JialianW/GRIT

0

相关内容

可理解性

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

脑心肌炎病毒受体的筛选、鉴定及功能分析

国家自然科学基金

0+阅读 · 2014年12月31日

基于PLL构建刺激响应性递送siRNA载体抗Her2阳性乳腺癌治疗研究

国家自然科学基金

0+阅读 · 2013年12月31日

KAT2B对脂肪细胞分化调控的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

15-kDa硒蛋白在内质网应激（ERS）和阿尔茨海默病(AD)中的功能研究

国家自然科学基金

0+阅读 · 2012年12月31日

肾脏单核-巨噬细胞系统中IKKα-p52:RelB途径活化对促进肾脏缺血再灌注损伤后修复的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

FRP加固钢筋混凝土框架填充墙结构的整体抗震性能试验、分析与设计方法

国家自然科学基金

0+阅读 · 2012年12月31日

结核分枝杆菌分泌脂蛋白对宿主巨噬细胞信号通路调控的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

斑马鱼心脏发育

国家自然科学基金

0+阅读 · 2009年12月31日

一个新的mRNA-like非编码RNA功能研究

国家自然科学基金

0+阅读 · 2008年12月31日

AT1受体激动抗体致动脉粥样硬化及其机制研究

国家自然科学基金

0+阅读 · 2008年12月31日

Leveraging Contaminated Datasets to Learn Clean-Data Distribution with Purified Generative Adversarial Networks

Arxiv

0+阅读 · 2023年2月3日

The Learnable Typewriter: A Generative Approach to Text Line Analysis

Arxiv

0+阅读 · 2023年2月3日

High-resolution Iterative Feedback Network for Camouflaged Object Detection

Arxiv

0+阅读 · 2023年2月3日

Aerial Image Object Detection With Vision Transformer Detector (ViTDet)

Arxiv

0+阅读 · 2023年2月2日

How to choose "Good" Samples for Text Data Augmentation

Arxiv

0+阅读 · 2023年2月2日

What Makes Good Examples for Visual In-Context Learning?

Arxiv

0+阅读 · 2023年2月1日

PV3D: A 3D Generative Model for Portrait Video Generation

Arxiv

0+阅读 · 2023年2月1日

Conditional Prompt Learning for Vision-Language Models

Conditional Prompt Learning for Vision-Language Models

Arxiv

13+阅读 · 2022年3月10日

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Arxiv

19+阅读 · 2020年2月15日

Transferring Common-Sense Knowledge for Object Detection

Arxiv

12+阅读 · 2018年4月3日

VIP会员

文章信息

相关主题

相关VIP内容

NeurlPS 2022 | 自然语言处理相关论文分类整理

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

165+阅读 · 2020年3月18日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

强化学习最新教程，17页pdf

强化学习最新教程，17页pdf

专知会员服务

182+阅读 · 2019年10月11日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

操作系统智能体：基于多模态大模型（MLLM）的通用计算设备智能体综述

《美国太空军系统全生命周期建模、仿真与分析效能提升方案》最新84页报告

【博士论文】推进数据高效的深度学习：非参数 Transformer、主动测试与上下文学习

自主人工智能：未来战争是否将是自主化的？

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium4

中国图象图形学学会CSIG

0+阅读 · 2021年11月10日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

相关论文

Leveraging Contaminated Datasets to Learn Clean-Data Distribution with Purified Generative Adversarial Networks

Arxiv

0+阅读 · 2023年2月3日

The Learnable Typewriter: A Generative Approach to Text Line Analysis

Arxiv

0+阅读 · 2023年2月3日

High-resolution Iterative Feedback Network for Camouflaged Object Detection

Arxiv

0+阅读 · 2023年2月3日

Aerial Image Object Detection With Vision Transformer Detector (ViTDet)

Arxiv

0+阅读 · 2023年2月2日

How to choose "Good" Samples for Text Data Augmentation

Arxiv

0+阅读 · 2023年2月2日

What Makes Good Examples for Visual In-Context Learning?

Arxiv

0+阅读 · 2023年2月1日

PV3D: A 3D Generative Model for Portrait Video Generation

Arxiv

0+阅读 · 2023年2月1日

Conditional Prompt Learning for Vision-Language Models

Conditional Prompt Learning for Vision-Language Models

Arxiv

13+阅读 · 2022年3月10日

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

UniViLM: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Arxiv

19+阅读 · 2020年2月15日

Transferring Common-Sense Knowledge for Object Detection

Arxiv

12+阅读 · 2018年4月3日

相关基金

脑心肌炎病毒受体的筛选、鉴定及功能分析

国家自然科学基金

0+阅读 · 2014年12月31日

基于PLL构建刺激响应性递送siRNA载体抗Her2阳性乳腺癌治疗研究

国家自然科学基金

0+阅读 · 2013年12月31日

KAT2B对脂肪细胞分化调控的分子机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

15-kDa硒蛋白在内质网应激（ERS）和阿尔茨海默病(AD)中的功能研究

国家自然科学基金

0+阅读 · 2012年12月31日

肾脏单核-巨噬细胞系统中IKKα-p52:RelB途径活化对促进肾脏缺血再灌注损伤后修复的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

FRP加固钢筋混凝土框架填充墙结构的整体抗震性能试验、分析与设计方法

国家自然科学基金

0+阅读 · 2012年12月31日

结核分枝杆菌分泌脂蛋白对宿主巨噬细胞信号通路调控的分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

斑马鱼心脏发育

国家自然科学基金

0+阅读 · 2009年12月31日

一个新的mRNA-like非编码RNA功能研究

国家自然科学基金

0+阅读 · 2008年12月31日

AT1受体激动抗体致动脉粥样硬化及其机制研究

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员