基于神经网络包装的夹心式视频压缩 (Sandwiched Video Compression: Efficiently Extending the Reach of Standard Codecs with Neural Wrappers) - 专知论文

会员服务 ·

0

HEVC · Networking · 近似 · 损失函数（机器学习） · Performer ·

2023 年 3 月 20 日

Sandwiched Video Compression: Efficiently Extending the Reach of Standard Codecs with Neural Wrappers

翻译：基于神经网络包装的夹心式视频压缩

Berivan Isik,Onur G. Guleryuz,Danhang Tang,Jonathan Taylor,Philip A. Chou

from arxiv, under review

We propose sandwiched video compression -- a video compression system that wraps neural networks around a standard video codec. The sandwich framework consists of a neural pre- and post-processor with a standard video codec between them. The networks are trained jointly to optimize a rate-distortion loss function with the goal of significantly improving over the standard codec in various compression scenarios. End-to-end training in this setting requires a differentiable proxy for the standard video codec, which incorporates temporal processing with motion compensation, inter/intra mode decisions, and in-loop filtering. We propose differentiable approximations to key video codec components and demonstrate that the neural codes of the sandwich lead to significantly better rate-distortion performance compared to compressing the original frames of the input video in two important scenarios. When transporting high-resolution video via low-resolution HEVC, the sandwich system obtains 6.5 dB improvements over standard HEVC. More importantly, using the well-known perceptual similarity metric, LPIPS, we observe $~30 \%$ improvements in rate at the same quality over HEVC. Last but not least we show that pre- and post-processors formed by very modestly-parameterized, light-weight networks can closely approximate these results.

翻译：我们提出了夹心式视频压缩，这是一种将神经网络封装在标准视频编解码器周围的视频压缩系统。夹层框架由一个神经网络前处理器和一个神经网络后处理器组成，在它们之间使用标准视频编解码器。这些网络是联合训练的，以优化速率-失真损失函数，旨在在各种压缩方案中显著改善标准编解码器的性能。在这种设置下的端到端训练需要一个可微分的标准视频编解码器代理，该代理集成了具有运动补偿、帧内和帧间模式决策以及循环滤波的时间处理。我们提出了关键视频编解码器组件的可微分近似，并证明了夹心代码比压缩原始输入视频帧在两个重要场景下显著提高了速率-失真性能。当通过低分辨率HEVC传输高分辨率视频时，夹心系统比标准HEVC获得了6.5 dB的改进。更重要的是，使用著名的感知相似度度量LPIPS，我们观察到在相同的质量下比HEVC提高了约30％的速率。最后但并非最不重要的是，我们展示了由非常适度参数化的轻量化网络形成的前处理器和后处理器可以紧密地近似这些结果。

0

相关内容

HEVC

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

专知会员服务

27+阅读 · 2022年3月3日

【论文推荐】逆问题，深度学习，对称性破缺，Inverse Problems, Deep Learning, and Symmetry Breaking

【论文推荐】逆问题，深度学习，对称性破缺，Inverse Problems, Deep Learning, and Symmetry Breaking

专知会员服务

26+阅读 · 2020年3月27日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【DeepMind】PolyGen: 一种三维网格的自回归生成模型，PolyGen: An Autoregressive Generative Model of 3D Meshes

【DeepMind】PolyGen: 一种三维网格的自回归生成模型，PolyGen: An Autoregressive Generative Model of 3D Meshes

专知会员服务

37+阅读 · 2020年2月27日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【NeurlPS2019论文总结】它是这样的:用于可解释图像识别的深度学习，This Looks Like That: Deep Learning for Interpretable Image Recognition

【NeurlPS2019论文总结】它是这样的:用于可解释图像识别的深度学习，This Looks Like That: Deep Learning for Interpretable Image Recognition

专知会员服务

22+阅读 · 2019年12月17日

【AAAI2020论文】小样本网络压缩，Few Shot Network Compression via Cross Distillation (附pdf）

专知会员服务

26+阅读 · 2019年11月23日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Deep Compression/Acceleration：模型压缩加速论文汇总

Deep Compression/Acceleration：模型压缩加速论文汇总

极市平台

14+阅读 · 2019年5月15日

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

泡泡机器人SLAM

23+阅读 · 2019年1月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【泡泡一分钟】基于李群的无损卡尔曼滤波器在视觉里程计上的应用

【泡泡一分钟】基于李群的无损卡尔曼滤波器在视觉里程计上的应用

泡泡机器人SLAM

11+阅读 · 2018年12月17日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

专知

10+阅读 · 2018年6月8日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

具有群作用CR流形上的Morse不等式

国家自然科学基金

0+阅读 · 2015年12月31日

Resveratrol联合MSCs移植对阿尔茨海默鼠的干预效果及Sirt1分子信号的介导作用

国家自然科学基金

0+阅读 · 2014年12月31日

面向DVS-MVI的多核SoC测试结构设计与测试调度算法协同优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

PD-L1和ICAM-1共修饰的MSCs诱导异体复合组织移植免疫耐受的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于HEVC的多视点视频加深度三维视频编码快速算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

复合壳层纳米电缆阵列的能带调控机制及其光伏特性

国家自然科学基金

0+阅读 · 2012年12月31日

PlncRNA-1在雄激素抵抗型前列腺癌中的作用及机制研究

国家自然科学基金

1+阅读 · 2012年12月31日

柔性η-CuPc纳米柱阵列有机薄膜太阳能电池

国家自然科学基金

0+阅读 · 2011年12月31日

基于动力学分析的Internet网络拥塞控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

约化群酉表示的branching law及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

MagicVideo: Efficient Video Generation With Latent Diffusion Models

Arxiv

0+阅读 · 2023年5月11日

V2Meow: Meowing to the Visual Beat via Music Generation

Arxiv

0+阅读 · 2023年5月11日

Treasure What You Have: Exploiting Similarity in Deep Neural Networks for Efficient Video Processing

Arxiv

0+阅读 · 2023年5月10日

Few-shot Action Recognition via Intra- and Inter-Video Information Maximization

Arxiv

0+阅读 · 2023年5月10日

Mobile Image Restoration via Prior Quantization

Arxiv

0+阅读 · 2023年5月10日

ToolCoder: Teach Code Generation Models to use API search tools

Arxiv

0+阅读 · 2023年5月9日

A Simple, Yet Effective Approach to Finding Biases in Code Generation

A Simple, Yet Effective Approach to Finding Biases in Code Generation

Arxiv

0+阅读 · 2023年5月9日

Measuring Forgetting of Memorized Training Examples

Arxiv

0+阅读 · 2023年5月9日

Training Graph Neural Networks with 1000 Layers

Arxiv

13+阅读 · 2021年6月14日

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Arxiv

13+阅读 · 2020年6月24日

VIP会员

文章信息

相关主题

损失函数（机器学习）

相关VIP内容

【2022新书】高效深度学习，Efficient Deep Learning Book

【2022新书】高效深度学习，Efficient Deep Learning Book

专知会员服务

125+阅读 · 2022年4月21日

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

【CVPR 2022】多模态视频字幕的端到端生成预训练，End-to-end Generative Pretraining for Multimodal Video Captioning

专知会员服务

27+阅读 · 2022年3月3日

【论文推荐】逆问题，深度学习，对称性破缺，Inverse Problems, Deep Learning, and Symmetry Breaking

【论文推荐】逆问题，深度学习，对称性破缺，Inverse Problems, Deep Learning, and Symmetry Breaking

专知会员服务

26+阅读 · 2020年3月27日

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

【SIGMOD2020-CMU】在内存中搜索树的顺序保持键压缩，Order-Preserving Key Compression for In-Memory Search Trees

专知会员服务

15+阅读 · 2020年3月7日

【DeepMind】PolyGen: 一种三维网格的自回归生成模型，PolyGen: An Autoregressive Generative Model of 3D Meshes

【DeepMind】PolyGen: 一种三维网格的自回归生成模型，PolyGen: An Autoregressive Generative Model of 3D Meshes

专知会员服务

37+阅读 · 2020年2月27日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【NeurlPS2019论文总结】它是这样的:用于可解释图像识别的深度学习，This Looks Like That: Deep Learning for Interpretable Image Recognition

【NeurlPS2019论文总结】它是这样的:用于可解释图像识别的深度学习，This Looks Like That: Deep Learning for Interpretable Image Recognition

专知会员服务

22+阅读 · 2019年12月17日

【AAAI2020论文】小样本网络压缩，Few Shot Network Compression via Cross Distillation (附pdf）

专知会员服务

26+阅读 · 2019年11月23日

《DeepGCNs: Making GCNs Go as Deep as CNNs》

《DeepGCNs: Making GCNs Go as Deep as CNNs》

专知会员服务

31+阅读 · 2019年10月17日

热门VIP内容

开通专知VIP会员享更多权益服务

【伯克利博士论文】通过真实世界实践赋能机器人自主性

军用无人机集群技术尚未成熟——但潜力可期

人工智能安全治理白皮书（2025）

AgentOps综述：分类、挑战与未来方向

相关资讯

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

Deep Compression/Acceleration：模型压缩加速论文汇总

Deep Compression/Acceleration：模型压缩加速论文汇总

极市平台

14+阅读 · 2019年5月15日

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

【泡泡一分钟】DS-SLAM: 动态环境下的语义视觉SLAM

泡泡机器人SLAM

23+阅读 · 2019年1月18日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

【泡泡一分钟】基于李群的无损卡尔曼滤波器在视觉里程计上的应用

【泡泡一分钟】基于李群的无损卡尔曼滤波器在视觉里程计上的应用

泡泡机器人SLAM

11+阅读 · 2018年12月17日

disentangled-representation-papers

disentangled-representation-papers

CreateAMind

26+阅读 · 2018年9月12日

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

【论文推荐】最新五篇视觉问答相关论文—深度学习评价、交互注意融合、VizWiz、引导注意力、

专知

10+阅读 · 2018年6月8日

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

ResNet, AlexNet, VGG, Inception：各种卷积网络架构的理解

全球人工智能

20+阅读 · 2017年12月17日

可解释的CNN

可解释的CNN

CreateAMind

17+阅读 · 2017年10月5日

相关论文

MagicVideo: Efficient Video Generation With Latent Diffusion Models

Arxiv

0+阅读 · 2023年5月11日

V2Meow: Meowing to the Visual Beat via Music Generation

Arxiv

0+阅读 · 2023年5月11日

Treasure What You Have: Exploiting Similarity in Deep Neural Networks for Efficient Video Processing

Arxiv

0+阅读 · 2023年5月10日

Few-shot Action Recognition via Intra- and Inter-Video Information Maximization

Arxiv

0+阅读 · 2023年5月10日

Mobile Image Restoration via Prior Quantization

Arxiv

0+阅读 · 2023年5月10日

ToolCoder: Teach Code Generation Models to use API search tools

Arxiv

0+阅读 · 2023年5月9日

A Simple, Yet Effective Approach to Finding Biases in Code Generation

A Simple, Yet Effective Approach to Finding Biases in Code Generation

Arxiv

0+阅读 · 2023年5月9日

Measuring Forgetting of Memorized Training Examples

Arxiv

0+阅读 · 2023年5月9日

Training Graph Neural Networks with 1000 Layers

Arxiv

13+阅读 · 2021年6月14日

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Minimal Variance Sampling with Provable Guarantees for Fast Training of Graph Neural Networks

Arxiv

13+阅读 · 2020年6月24日

相关基金

具有群作用CR流形上的Morse不等式

国家自然科学基金

0+阅读 · 2015年12月31日

Resveratrol联合MSCs移植对阿尔茨海默鼠的干预效果及Sirt1分子信号的介导作用

国家自然科学基金

0+阅读 · 2014年12月31日

面向DVS-MVI的多核SoC测试结构设计与测试调度算法协同优化方法研究

国家自然科学基金

0+阅读 · 2013年12月31日

PD-L1和ICAM-1共修饰的MSCs诱导异体复合组织移植免疫耐受的机制研究

国家自然科学基金

0+阅读 · 2013年12月31日

基于HEVC的多视点视频加深度三维视频编码快速算法研究

国家自然科学基金

0+阅读 · 2013年12月31日

复合壳层纳米电缆阵列的能带调控机制及其光伏特性

国家自然科学基金

0+阅读 · 2012年12月31日

PlncRNA-1在雄激素抵抗型前列腺癌中的作用及机制研究

国家自然科学基金

1+阅读 · 2012年12月31日

柔性η-CuPc纳米柱阵列有机薄膜太阳能电池

国家自然科学基金

0+阅读 · 2011年12月31日

基于动力学分析的Internet网络拥塞控制研究

国家自然科学基金

0+阅读 · 2009年12月31日

约化群酉表示的branching law及其应用

国家自然科学基金

0+阅读 · 2009年12月31日

微信扫码咨询专知VIP会员