使用 cINNs 的Stochatic 图像合成合成 (Stochastic Image-to-Video Synthesis using cINNs)

Video understanding calls for a model to learn the characteristic interplay between static scene content and its dynamics: Given an image, the model must be able to predict a future progression of the portrayed scene and, conversely, a video should be explained in terms of its static image content and all the remaining characteristics not present in the initial frame. This naturally suggests a bijective mapping between the video domain and the static content as well as residual information. In contrast to common stochastic image-to-video synthesis, such a model does not merely generate arbitrary videos progressing the initial image. Given this image, it rather provides a one-to-one mapping between the residual vectors and the video with stochastic outcomes when sampling. The approach is naturally implemented using a conditional invertible neural network (cINN) that can explain videos by independently modelling static and other video characteristics, thus laying the basis for controlled video synthesis. Experiments on four diverse video datasets demonstrate the effectiveness of our approach in terms of both the quality and diversity of the synthesized results. Our project page is available at https://bit.ly/3t66bnU.

翻译：视频理解要求一种模型来了解静态现场内容及其动态之间的特征相互作用:根据图像,模型必须能够预测描绘的场景的未来进展,反之,应当用其静态图像内容和初始框架没有显示的所有其余特征来解释视频。这自然意味着视频域与静态内容以及剩余信息之间的双向映射。与普通的随机图像合成相比,这种模型不仅仅是在初始图像上产生任意的视频。鉴于这一图像,它反而提供了残余矢量和视频之间的一对一映图,在取样时带有随机结果。该方法自然使用一个有条件的不可逆神经网络(cN)来解释视频,通过独立模拟静态和其他视频特征来解释视频,从而为受控的视频合成奠定基础。在四个不同的视频数据集上进行的实验表明我们的方法在综合结果的质量和多样性方面的有效性。我们的项目网页可在 https://bit.ly/366bnU上查阅。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/