DPTNet: 用于现场文字探测的双纸式变换结构 (DPTNet: A Dual-Path Transformer Architecture for Scene Text Detection) - 专知论文

会员服务 ·

0

Attention · CLUES · Extensibility · 变换 · 卷积 ·

2022 年 8 月 21 日

DPTNet: A Dual-Path Transformer Architecture for Scene Text Detection

翻译：DPTNet: 用于现场文字探测的双纸式变换结构

Jingyu Lin,Jie Jiang,Yan Yan,Chunchao Guo,Hongfa Wang,Wei Liu,Hanzi Wang

The prosperity of deep learning contributes to the rapid progress in scene text detection. Among all the methods with convolutional networks, segmentation-based ones have drawn extensive attention due to their superiority in detecting text instances of arbitrary shapes and extreme aspect ratios. However, the bottom-up methods are limited to the performance of their segmentation models. In this paper, we propose DPTNet (Dual-Path Transformer Network), a simple yet effective architecture to model the global and local information for the scene text detection task. We further propose a parallel design that integrates the convolutional network with a powerful self-attention mechanism to provide complementary clues between the attention path and convolutional path. Moreover, a bi-directional interaction module across the two paths is developed to provide complementary clues in the channel and spatial dimensions. We also upgrade the concentration operation by adding an extra multi-head attention layer to it. Our DPTNet achieves state-of-the-art results on the MSRA-TD500 dataset, and provides competitive results on other standard benchmarks in terms of both detection accuracy and speed.

翻译：深层学习的繁荣有助于现场文字探测的迅速进展。在与革命网络的所有方法中,以分裂为基础的方法已引起广泛的注意,因为它们在探测任意形状和极端方面比率的文字实例方面具有优越性。然而,自下而上的方法仅限于其分解模型的性能。在本文件中,我们提议DPTNet(Dual-Path变异器网络),这是一个简单而有效的结构,用以模拟现场文字探测任务的全球和地方信息。我们进一步提议了一种平行的设计,将革命网络与一个强大的自我注意机制结合起来,以提供注意力路径和相联路径之间的互补线索。此外,正在开发一个双向互动模块,以提供通道和空间层面的互补线索。我们还通过增加一个多头关注层来提升集中操作。我们的DPTNet在MSRA-TD500数据集上取得了最先进的结果,并在探测准确性和速度方面提供其他标准基准的竞争性结果。

0

相关内容

Attention

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

【推荐】全卷积语义分割综述

【推荐】全卷积语义分割综述

机器学习研究会

19+阅读 · 2017年8月31日

肺炎支原体外排泵ABC Transporter在大环内酯类耐药中的作用机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

小尺度电离层扰动的TEC起伏特性研究

国家自然科学基金

0+阅读 · 2014年12月31日

褐家鼠群体MHC遗传多态性与SEOV感染相关性的研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于Archimedean三角模的区间犹豫模糊平均型集结算子及其在决策中的应用

国家自然科学基金

0+阅读 · 2012年12月31日

统一贝叶斯框架下的活动轮廓模型OCT心管图像序列分割

国家自然科学基金

0+阅读 · 2012年12月31日

面向QoS需求的智能电网邻域网中多业务信息接入控制

国家自然科学基金

0+阅读 · 2012年12月31日

虫草多糖活性片段的筛选及分子作用机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

行为决策视角下考虑产品生命周期的数字盗版控制策略组合研究

国家自然科学基金

0+阅读 · 2012年12月31日

神经信息流与突触传递可塑性

国家自然科学基金

0+阅读 · 2011年12月31日

基于遥感的棉花长势监测模型及其栽培应用

国家自然科学基金

0+阅读 · 2008年12月31日

MGTR: End-to-End Mutual Gaze Detection with Transformer

Arxiv

0+阅读 · 2022年10月6日

Centralized Feature Pyramid for Object Detection

Arxiv

0+阅读 · 2022年10月5日

Perspective Aware Road Obstacle Detection

Arxiv

0+阅读 · 2022年10月4日

Cross-Modality Fusion Transformer for Multispectral Object Detection

Arxiv

0+阅读 · 2022年10月4日

An efficient encoder-decoder architecture with top-down attention for speech separation

Arxiv

0+阅读 · 2022年9月30日

Text Detection and Recognition in the Wild: A Review

Arxiv

20+阅读 · 2020年6月8日

Hierarchical Graph Pooling with Structure Learning

Arxiv

13+阅读 · 2019年11月14日

Reverse Attention for Salient Object Detection

Arxiv

11+阅读 · 2019年4月15日

Rotation-Sensitive Regression for Oriented Scene Text Detection

Arxiv

13+阅读 · 2018年3月14日

DOTA: A Large-scale Dataset for Object Detection in Aerial Images

Arxiv

19+阅读 · 2018年1月27日

VIP会员

文章信息

相关主题

相关VIP内容

50+篇《神经架构搜索NAS》2020论文合集

专知会员服务

61+阅读 · 2020年3月19日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Stabilizing Transformers for Reinforcement Learning

Stabilizing Transformers for Reinforcement Learning

专知会员服务

60+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用

专知会员服务

41+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

最新《扩散模型原理》新书，470页pdf

无人机作战：演进、创新与未来战场

AI 智能体简史

多模态空间推理在大模型时代：综述与基准测试

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

无监督元学习表示学习

无监督元学习表示学习

CreateAMind

27+阅读 · 2019年1月4日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

【推荐】ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

机器学习研究会

20+阅读 · 2017年12月17日

【推荐】深度学习目标检测概览

【推荐】深度学习目标检测概览

机器学习研究会

10+阅读 · 2017年9月1日

【推荐】全卷积语义分割综述

【推荐】全卷积语义分割综述

机器学习研究会

19+阅读 · 2017年8月31日

相关论文

MGTR: End-to-End Mutual Gaze Detection with Transformer

Arxiv

0+阅读 · 2022年10月6日

Centralized Feature Pyramid for Object Detection

Arxiv

0+阅读 · 2022年10月5日

Perspective Aware Road Obstacle Detection

Arxiv

0+阅读 · 2022年10月4日

Cross-Modality Fusion Transformer for Multispectral Object Detection

Arxiv

0+阅读 · 2022年10月4日

An efficient encoder-decoder architecture with top-down attention for speech separation

Arxiv

0+阅读 · 2022年9月30日

Text Detection and Recognition in the Wild: A Review

Arxiv

20+阅读 · 2020年6月8日

Hierarchical Graph Pooling with Structure Learning

Arxiv

13+阅读 · 2019年11月14日

Reverse Attention for Salient Object Detection

Arxiv

11+阅读 · 2019年4月15日

Rotation-Sensitive Regression for Oriented Scene Text Detection

Arxiv

13+阅读 · 2018年3月14日

DOTA: A Large-scale Dataset for Object Detection in Aerial Images

Arxiv

19+阅读 · 2018年1月27日

相关基金

肺炎支原体外排泵ABC Transporter在大环内酯类耐药中的作用机制研究

国家自然科学基金

0+阅读 · 2016年12月31日

小尺度电离层扰动的TEC起伏特性研究

国家自然科学基金

0+阅读 · 2014年12月31日

褐家鼠群体MHC遗传多态性与SEOV感染相关性的研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于Archimedean三角模的区间犹豫模糊平均型集结算子及其在决策中的应用

国家自然科学基金

0+阅读 · 2012年12月31日

统一贝叶斯框架下的活动轮廓模型OCT心管图像序列分割

国家自然科学基金

0+阅读 · 2012年12月31日

面向QoS需求的智能电网邻域网中多业务信息接入控制

国家自然科学基金

0+阅读 · 2012年12月31日

虫草多糖活性片段的筛选及分子作用机制的研究

国家自然科学基金

0+阅读 · 2012年12月31日

行为决策视角下考虑产品生命周期的数字盗版控制策略组合研究

国家自然科学基金

0+阅读 · 2012年12月31日

神经信息流与突触传递可塑性

国家自然科学基金

0+阅读 · 2011年12月31日

基于遥感的棉花长势监测模型及其栽培应用

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员