语言相关行动股和基于多模式代表组合的音频驱动的有声人一代 (Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion) - 专知论文

会员服务 ·

0

多峰值 · INFORMS · INTERACT · 相关系数 · 模型评估 ·

2022 年 4 月 27 日

Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusion

翻译：语言相关行动股和基于多模式代表组合的音频驱动的有声人一代

Sen Chen,Zhilei Liu,Jiaxing Liu,Longbiao Wang

from arxiv, arXiv admin note: text overlap with arXiv:2110.09951

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information, but also the local driving information of the mouth muscles. In this study, we propose a novel generative framework that contains a dilated non-causal temporal convolutional self-attention network as a multimodal fusion module to promote the relationship learning of cross-modal features. In addition, our proposed method uses both audio- and speech-related facial action units (AUs) as driving information. Speech-related AU information can guide mouth movements more accurately. Because speech is highly correlated with speech-related AUs, we propose an audio-to-AU module to predict speech-related AU information. We utilize pre-trained AU classifier to ensure that the generated images contain correct AU information. We verify the effectiveness of the proposed model on the GRID and TCD-TIMIT datasets. An ablation study is also conducted to verify the contribution of each component. The results of quantitative and qualitative experiments demonstrate that our method outperforms existing methods in terms of both image quality and lip-sync accuracy.

翻译：在这项研究中,我们提出了一个新的基因框架,其中包含一个非因果的超时演进自留网络,作为多式组合模块,以促进跨模式特征的关系学习。此外,我们提议的方法同时使用音频和与语音有关的面部动作单位(AUS)作为驱动信息的工具。与发言有关的AU信息可以更准确地指导口腔运动。由于语音与AU高度相关,我们提议了一个音频到AU模块,以预测与AU有关的声音信息。我们使用事先培训的AU分类器,以确保生成的图像包含正确的AU信息。我们核查了全球资源数据库和TCD-TIMIT数据集的拟议模型的有效性。还进行了一项模拟研究,以核实每个组件的贡献。定量和定性实验的结果表明,我们的方法在质量和图像方面都超过了现有方法的准确性。

0

相关内容

多峰值

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

CSP-GNPs对宫颈癌的靶向放疗增敏作用及其机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

石墨烯量子点-银异质结表面等离激元共振增强共轭聚合物发光研究

国家自然科学基金

0+阅读 · 2013年12月31日

磁场诱导CNTs有序排列的高分子-无机杂化膜的制备

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

雌激素受体alpha亚基118位点丝氨酸的磷酸化在大脑海马性别分化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

低氧诱导因子1α（HIF-1α）的SUMO化修饰调控对髓核细胞氧张力耐受的影响及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Ghrelin对胰岛β细胞分泌胰岛素和增殖的影响及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

Tecto调节非洲爪蛙胚层决定与分化的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预G蛋白偶联受体40对妊娠期糖尿病大鼠胰岛素抵抗及糖稳态的影响

国家自然科学基金

0+阅读 · 2012年12月31日

磁性铁电材料的第一原理研究与大尺度相结构模拟

国家自然科学基金

0+阅读 · 2011年12月31日

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

Arxiv

0+阅读 · 2022年6月15日

Learning to Reduce Information Bottleneck for Object Detection in Aerial Images

Arxiv

0+阅读 · 2022年6月15日

Mitigating Bias in Facial Analysis Systems by Incorporating Label Diversity

Mitigating Bias in Facial Analysis Systems by Incorporating Label Diversity

Arxiv

0+阅读 · 2022年6月14日

SelfReformer: Self-Refined Network with Transformer for Salient Object Detection

Arxiv

0+阅读 · 2022年6月14日

Multispectral image fusion by super pixel statistics

Arxiv

0+阅读 · 2022年6月12日

Head and eye egocentric gesture recognition for human-robot interaction using eyewear cameras

Arxiv

0+阅读 · 2022年6月10日

A Survey on Neural Speech Synthesis

Arxiv

14+阅读 · 2021年6月30日

Affective Image Content Analysis: Two Decades Review and New Perspectives

Arxiv

16+阅读 · 2021年6月30日

MVFNet: Multi-View Fusion Network for Efficient Video Recognition

Arxiv

13+阅读 · 2021年1月5日

Aspect Based Sentiment Analysis with Gated Convolutional Networks

Arxiv

12+阅读 · 2018年5月18日

VIP会员

文章信息

相关主题

相关VIP内容

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

CVPR 2020 论文开源项目合集

专知会员服务

110+阅读 · 2020年3月12日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

Aspect-Oriented Syntax Network for Aspect-Based Sentiment Analysis，中山大学数据科学与计算机学院权小军教授，第八届全国社会媒体处理大会SMP2019

专知会员服务

19+阅读 · 2019年10月22日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

[综述]深度学习下的场景文本检测与识别

[综述]深度学习下的场景文本检测与识别

专知会员服务

78+阅读 · 2019年10月10日

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

【CMU卡内基梅隆大学】深度学习在计算机视觉的应用：方法，解释，因果与公平性

专知会员服务

83+阅读 · 2019年10月9日

最新BERT相关论文清单，BERT-related Papers

最新BERT相关论文清单，BERT-related Papers

专知会员服务

53+阅读 · 2019年9月29日

热门VIP内容

开通专知VIP会员享更多权益服务

【博士论文】在低维和高维空间中分析、建模和转换潜在表征

从无人机到数据：揭示边缘计算作为新作战域

可解释人工智能的基础

大规模视觉模型中的基于提示的适应：综述

相关资讯

VCIP 2022 Call for Special Session Proposals

VCIP 2022 Call for Special Session Proposals

CCF多媒体专委会

1+阅读 · 2022年4月1日

IEEE ICKG 2022: Call for Papers

IEEE ICKG 2022: Call for Papers

机器学习与推荐算法

3+阅读 · 2022年3月30日

ACM MM 2022 Call for Papers

ACM MM 2022 Call for Papers

CCF多媒体专委会

5+阅读 · 2022年3月29日

AIART 2022 Call for Papers

AIART 2022 Call for Papers

CCF多媒体专委会

1+阅读 · 2022年2月13日

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

【ICIG2021】Check out the hot new trailer of ICIG2021 Symposium8

中国图象图形学学会CSIG

0+阅读 · 2021年11月16日

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

会议交流 | IJCKG: International Joint Conference on Knowledge Graphs

开放知识图谱

0+阅读 · 2021年9月9日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

【论文推荐】最新八篇情感分析相关论文—Pair-wise判别器、多模态情感分析、上下文语境、Gated 卷积网络

专知

20+阅读 · 2018年6月29日

相关论文

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

Arxiv

0+阅读 · 2022年6月15日

Learning to Reduce Information Bottleneck for Object Detection in Aerial Images

Arxiv

0+阅读 · 2022年6月15日

Mitigating Bias in Facial Analysis Systems by Incorporating Label Diversity

Mitigating Bias in Facial Analysis Systems by Incorporating Label Diversity

Arxiv

0+阅读 · 2022年6月14日

SelfReformer: Self-Refined Network with Transformer for Salient Object Detection

Arxiv

0+阅读 · 2022年6月14日

Multispectral image fusion by super pixel statistics

Arxiv

0+阅读 · 2022年6月12日

Head and eye egocentric gesture recognition for human-robot interaction using eyewear cameras

Arxiv

0+阅读 · 2022年6月10日

A Survey on Neural Speech Synthesis

Arxiv

14+阅读 · 2021年6月30日

Affective Image Content Analysis: Two Decades Review and New Perspectives

Arxiv

16+阅读 · 2021年6月30日

MVFNet: Multi-View Fusion Network for Efficient Video Recognition

Arxiv

13+阅读 · 2021年1月5日

Aspect Based Sentiment Analysis with Gated Convolutional Networks

Arxiv

12+阅读 · 2018年5月18日

相关基金

CSP-GNPs对宫颈癌的靶向放疗增敏作用及其机制的研究

国家自然科学基金

0+阅读 · 2014年12月31日

石墨烯量子点-银异质结表面等离激元共振增强共轭聚合物发光研究

国家自然科学基金

0+阅读 · 2013年12月31日

磁场诱导CNTs有序排列的高分子-无机杂化膜的制备

国家自然科学基金

0+阅读 · 2013年12月31日

Intraflagellar Transport运输纤毛蛋白的分子机理

国家自然科学基金

0+阅读 · 2012年12月31日

雌激素受体alpha亚基118位点丝氨酸的磷酸化在大脑海马性别分化中的作用及机制

国家自然科学基金

0+阅读 · 2012年12月31日

低氧诱导因子1α（HIF-1α）的SUMO化修饰调控对髓核细胞氧张力耐受的影响及机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

Ghrelin对胰岛β细胞分泌胰岛素和增殖的影响及分子机制

国家自然科学基金

0+阅读 · 2012年12月31日

Tecto调节非洲爪蛙胚层决定与分化的机制研究

国家自然科学基金

0+阅读 · 2012年12月31日

靶向干预G蛋白偶联受体40对妊娠期糖尿病大鼠胰岛素抵抗及糖稳态的影响

国家自然科学基金

0+阅读 · 2012年12月31日

磁性铁电材料的第一原理研究与大尺度相结构模拟

国家自然科学基金

0+阅读 · 2011年12月31日

微信扫码咨询专知VIP会员