ABOUT

Ye Liu.

Ye Liu

I’m a Research Scientist working on Multi-modal Agents. I received my PhD from The Hong Kong Polytechnic University, advised by Prof. Chang Wen Chen. Before that, I obtained my B.S. & B.E. from Wuhan University. I’m fortunate to work at ByteDance Seed, Tencent ARC Lab, Show Lab @ NUS, VCG @ Harvard, SUNY Buffalo, and CUHK-Shenzhen during my research journey.

My research lies in building next-generation multi-modal agents that can perceive, reason, and interact with the digital and physical world:

Please feel free to reach out if you are interested in related topics :)

News

  1. I received PolyU COMP Outstanding PhD Thesis Award.

  2. Our Ground3D-LMM got accepted by ECCV 2026.

  3. I'm awarded RGC Junior Research Fellow Scheme (JRFS).

  4. Check out my work in PolyU Top 10 Research & Innovation Stories of the Year.

  5. Our VideoMind got accepted by ICLR 2026.

More ↓Less ↑
  1. I received PolyU Distinguished Postdoctoral Fellowship.

  2. I received NeurIPS Scholar Award.

  3. Our VTG LLM Survey got accepted by TPAMI.

  4. Our VideoMind got accepted by LAW @ NeurIPS 2025 (Spotlight).

  5. Our new work UniPixel got accepted by NeurIPS 2025.

  6. One paper got accepted by ICCV 2025.

  7. One paper got accepted by NeurIPS 2024.

  8. One paper and its demo got accepted by ECCV 2024.

  9. I'm joining Show Lab @ NUS as a visiting student.

  10. One paper got accepted by TNNLS.

  11. One paper got accepted by CIKM 2023.

  12. I will join Visual Computing Group @ Harvard as an Associate.

  13. My startup on AI + Healthcare is granted by HKSTP for ideation.

  14. I received The Most Appreciated Teaching Assistant (MATA) award.

  15. I received PolyU Student Innovation and Entrepreneurship Scholarship.

  16. One paper got accepted by CVPR 2022.

  17. I started my internship at ARC Lab, Tencent PCG.

  18. One paper got accepted by ACM Multimedia 2020.

  19. I graduated from Wuhan University with honors.

  20. I started my internship at SUNY-Buffalo and CUHK-Shenzhen.

Publications

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning — overview

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

Ye Liu*, Kevin Qinghong Lin*, Chang Wen Chen, Mike Zheng Shou

International Conference on Learning Representations (ICLR) · 2026
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM — overview

Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

Amol Harsh, Zongyan Han, Jean Lahoud, Ye Liu, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Shahbaz Khan

The European Conference on Computer Vision (ECCV) · 2026
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning — overview

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

Ye Liu, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu, Ying Shan, Chang Wen Chen

Advances in Neural Information Processing Systems (NeurIPS) · 2025
A Survey on Video Temporal Grounding with Multimodal Large Language Model — overview

A Survey on Video Temporal Grounding with Multimodal Large Language Model

Jianlong Wu, Wei Liu, Ye Liu, Meng Liu, Liqiang Nie, Zhouchen Lin, Chang Wen Chen

IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) · 2025
VisionMath: Vision-Form Mathematical Problem-Solving — overview

VisionMath: Vision-Form Mathematical Problem-Solving

Zongyang Ma, Yuxin Chen, Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Shaojie Zhu, Chengxiang Zhuo, Bing Li, Ye Liu, Zang Li, Ying Shan, Weiming Hu

The IEEE/CVF International Conference on Computer Vision (ICCV) · 2025
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion — overview

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

Jixuan He*, Wanhua Li*, Ye Liu, Junsik Kim, Donglai Wei, Hanspeter Pfister

P13N: Personalization in Generative AI @ ICCV · 2025
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding — overview

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Ye Liu, Zongyang Ma, Zhongang Qi, Yang Wu, Ying Shan, Chang Wen Chen

Advances in Neural Information Processing Systems (NeurIPS) · 2024
R²-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding — overview

R²-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

Ye Liu, Jixuan He, Wanhua Li, Junsik Kim, Donglai Wei, Hanspeter Pfister, Chang Wen Chen

The European Conference on Computer Vision (ECCV) · 2024
Learning to Aggregate Multi-Scale Context for Instance Segmentation in Remote Sensing Images — overview

Learning to Aggregate Multi-Scale Context for Instance Segmentation in Remote Sensing Images

Ye Liu, Huifang Li, Chao Hu, Shuang Luo, Yan Luo, Chang Wen Chen

IEEE Transactions on Neural Networks and Learning Systems (TNNLS) · 2024
Timestamps as Prompts for Geography-Aware Location Recommendation — overview

Timestamps as Prompts for Geography-Aware Location Recommendation

Yan Luo, Haoyi Duan, Ye Liu, Fu-lai Chung

The ACM International Conference on Information and Knowledge Management (CIKM) · 2023
End-to-End Personalized Next Location Recommendation via Contrastive User Preference Modeling — overview

End-to-End Personalized Next Location Recommendation via Contrastive User Preference Modeling

Yan Luo, Ye Liu, Fu-lai Chung, Yu Liu, Takahiro Yabe

IEEE Transactions on Computational Social Systems (TCSS) · 2023
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection — overview

UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

Ye Liu, Siyuan Li, Yang Wu, Chang Wen Chen, Ying Shan, Xiaohu Qie

The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) · 2022
ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction Detection — overview

ConsNet: Learning Consistency Graph for Zero-Shot Human-Object Interaction Detection

Ye Liu, Junsong Yuan, Chang Wen Chen

The ACM International Conference on Multimedia (ACM MM) · 2020