【VALSE论文速览-223期】Scene Graph Disentanglement and Composition for Generalizable...

765
0
2025-04-14 22:06:10
5
投币
6
3
论文题目:Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation(NeurIPS 2024 Spotlight) 作者列表:Yunnan Wang, Ziqiang Li, Zequn Zhang, Wenyao Zhang, Baao Xie, Xihui Liu, Wenjun Zeng, Xin Jin 原文链接:https://arxiv.org/abs/2410.00447 论文摘要:There has been exciting progress in generating images from natural language or layout conditions. However, these methods struggle to faithfully reproduce complex scenes due to the insufficient modeling of multiple objects and their relationships. To address this issue, we leverage the scene graph, a powerful structured representation, for complex image generation. Different from the previous works that directly use scene graphs for generation, we employ the generative capabilities of variational autoencoders and diffusion models in a generalizable manner, compositing diverse disentangled visual clues from scene graphs. Specifically, we first propose a Semantics-Layout Variational AutoEncoder (SL-VAE) to jointly derive (layouts, semantics) from the input scene graph, which allows a more diverse and reasonable generation in a one-to-many mapping. We then develop a Compositional Masked Attention (CMA) integrated with a diffusion model, incorporating (layouts, semantics) with fine-grained attributes as generation guidance. To further achieve graph manipulation while keeping the visual content consistent, we introduce a Multi-Layered Sampler (MLS) for an "isolated" image editing effect. Extensive experiments demonstrate that our method outperforms recent competitors based on text, layout, or scene graph, in terms of generation rationality and controllability. 视频讲者简介:王允楠,上海交通大学和宁波东方理工大学联合培养博士,研究兴趣包括复杂场景下的生成与理解等,在NeurlPS、IJCAI等AI顶级会议上发表多篇论文。
为计算机视觉、图像处理、模式识别与机器学习等研究领域内的华人青年学者提供深入学术交流的舞台。
客服
顶部
赛事库 课堂 2021拜年纪