Publications

* Equal contribution, ✉ Corresponding author

2026

  1. milr_m.png
    MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
    Yapeng Mi , Yanpeng Zhao✉ , Hengli Li , Chenxi Li , Huimin Wu , Xiaojian Ma , Song-Chun Zhu , Ying Nian Wu , and Qing Li✉
    International Conference on Learning Representations (ICLR), 2026
  2. scenedreamer360.jpg
    SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
    Wenrui Li , Fucheng Cai , Yapeng Mi , Zhe Yang , Wangmeng Zuo , Xingtao Wang , and Xiaopeng Fan✉
    IEEE Transactions on Multimedia, 2026

2025

  1. sport.png
    Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
    Pengxiang Li* , Zhi Gao* , Bofei Zhang , Yapeng Mi , Xiaojian Ma , Chenrui Shi , Tao Yuan , Yuwei Wu✉ , Yunde Jia , Song-Chun Zhu , and Qing Li✉
    Advances in Neural Information Processing Systems (NeurIPS), 2025
  2. building.png
    Building LLM Agents by Incorporating Insights from Computer Systems
    Yapeng Mi , Zhi Gao , Xiaojian Ma , and Qing Li✉
    arXiv preprint arXiv:2504.04485, 2025