Shangzhe Di(狄尚哲)

PhD Candidate at SJTU · Student Researcher at Google DeepMind

shangzhe.di@gmail.com / CV / Google Scholar / GitHub

Hi, I am a fourth-year PhD candidate at Shanghai Jiao Tong University, advised by Prof. Weidi Xie. My research focuses on video understanding and multimodal learning.

I am currently a Student Researcher at Google DeepMind in London, working on agentic 4D generation. Previously, I interned at ByteDance Seed and Alibaba.

I'm graduating in 2027 and open to opportunities in both industry and academia. Feel free to reach out if you think there's a good fit.

Shangzhe Di

Experience

Google DeepMindMay 2026 – Nov 2026
Student Researcher · Agentic 4D Generation
ByteDance SeedNov 2024 – Dec 2025
TopSeed Intern · Visual Representation Learning
AlibabaApr 2024 – Sep 2024
Research Intern · Streaming Video Understanding

Education

Shanghai Jiao Tong University2023 – now
PhD · advised by Prof. Weidi Xie
Beihang University2020 – 2023
M.Eng. Computer Science · advised by Prof. Si Liu
Beihang University2016 – 2020
B.Eng. Software Engineering

Publications * equal contribution

2026
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
Yibin Yan, Jilan Xu, Shangzhe Di, Haoning Wu, Weidi Xie
ECCV, 2026

A unified streaming visual backbone for perception, reconstruction, and action.

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation
Shangzhe Di, Zhaokai Wang, Weidi Xie
BMVC, 2026

Image generators exhibit strong zero-shot perception capabilities through text prompting.

Revisiting Multi-Task Visual Representation Learning
Shangzhe Di, Zhonghua Zhai, Weidi Xie
ACCV, 2026

Multitask pre-training can yield scalable general-purpose visual representations.

2025
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
Zeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang, Yanfeng Wang, Weidi Xie
NeurIPS, 2025

Towards universal video grounding with superior accuracy, generalizability, and robustness.

Learning Streaming Video Representation via Multitask Training
Yibin Yan*, Jilan Xu*, Shangzhe Di, Yikun Liu, Yudi Shi, Qirui Chen, Zeqian Li, Yifei Huang, Weidi Xie
ICCV, 2025 Oral

Streaming video representations at global, temporal and spatial granularity, learned via multitask training.

Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
Yudi Shi, Shangzhe Di, Qirui Chen, Weidi Xie
CVPR, 2025

Distill multi-step reasoning and spatial-temporal understanding into a generative Video-LLM.

Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
Shangzhe Di, Zhelun Yu, et al.
ICLR, 2025

A training-free approach enabling Video-LLMs for streaming video question-answering.

Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
Qirui Chen, Shangzhe Di, Weidi Xie
AAAI, 2025

Pinpoint scattered visual evidence in long egocentric videos while responding to questions.

2024
Grounded Question-Answering in Long Egocentric Videos
Shangzhe Di, Weidi Xie
CVPR, 2024

Simultaneous query grounding and answering in long, egocentric videos.

2023
Linker: Learning Long Short-term Associations for Robust Visual Tracking
Zizheng Xun*, Shangzhe Di*, Yulu Gao, Zongheng Tang, Gang Wang, Si Liu, Bo Li
IEEE Transactions on Multimedia (TMM), 2023

Associate the initial template with a fast-updated reference region for robust visual tracking.

2021
Video Background Music Generation with Controllable Music Transformer
Shangzhe Di*, Zeren Jiang*, Si Liu, Zhaokai Wang, Leyan Zhu, Zexin He, Hongming Liu, Shuicheng Yan
ACM MM, 2021 Best Paper Award

The first satisfying method for video background music generation.

Honors & Awards

Academic Service

Reviewer of CVPR, ICCV, ECCV, ICLR, NeurIPS, and ICML.