【VALSE论文速览-154期】mPLUG-2: A Modularized Multi-modal Foundation Model Across Text…

1159
0
2023-12-01 17:02:53
18
9
12
7
论文题目:mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 论文出处:ICML 2023 论文摘要:近年来,语言、视觉和多模态预训练发生了巨大的融合。在本文中,我们提出了mPLUG-2,这是一种新的统一范例,具有模块化设计,可从模态协作中受益,同时解决模态纠缠的问题。与仅依赖序列到序列生成或基于编码器的实例鉴别的主导范例不同,mPLUG-2引入了一个多模块组合网络,通过共享通用的通用模块实现模态协作,并对不同的模态模块进行解缠以处理模态纠缠。它可以灵活选择不同的模块以处理所有模态的不同理解和生成任务,包括文本、图像和视频。实证研究表明,mPLUG-2在超过30个下游任务的广泛范围内实现了最先进或具有竞争力的结果,涵盖了图像-文本和视频-文本理解和生成的多模态任务,以及仅文本、仅图像和仅视频的单模态任务。值得注意的是,mPLUG-2在具有更小的模型大小和数据规模的情况下,在具有挑战性的MSRVTT视频QA和视频描述任务上展示了48.0的top-1准确度和80.3的CIDEr的新的最先进结果。它还展示了在视觉语言和视频语言任务上的强大的零样本可转移性。 作者列表:Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou (阿里巴巴达摩院) 讲者简介: Qinghao Ye is an Algorithm Engineer in DAMO Academy, Alibaba Group. He received the M.S. degree in Computer Scient & Engineering from University of California, San Diego in 2022. He has authored multiple publications, including works represented on ICCV, ACL and ICML. His work has been cited over 400 times according to Google Scholar, and has an H-index of 9. 参考文献: [1]  Xu H, Ye Q, Yan M, et al. mplug-2: A modularized multi-modal foundation model across text, image and video[C]. ICML 2023. [2] Ye Q, Xu H, Xu G, et al. mplug-owl: Modularization empowers large language models with multimodality[J]. arXiv preprint arXiv:2304.14178, 2023.
为计算机视觉、图像处理、模式识别与机器学习等研究领域内的华人青年学者提供深入学术交流的舞台。
客服
顶部
赛事库 课堂 2021拜年纪