1.陈铠 香港科技大学博士生 EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions 2.王奉祥 国防科技大学博士生 XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
3.洪文逸 清华大学博士生 MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
4.张天舒 清华大学本科生 ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
5.行习铭 北京航空航天大学博士生 Empowering LLMs to Understand and Generate Complex Vector Graphics
6.陈天宇 北京航空航天大学博士生 Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
7.闵安娜 清华大学本科生 Supervising Sound Localization using In-the-wild Ego-motion
8.张泽锋 中科院信息工程研究所博士生 Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
9.杜恒辉 中国人民大学硕士生 Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
10.梁思源 新加坡国立大学博士后研究员 Revisiting Backdoor Attacks against Large Vision-Language Models from Domain Shift