|
Shilin Ma
Hi! I am a PhD student at
Tsinghua University,
advised by Prof. Yansong Tang.
Previously, I received my B.Eng. degree in
Electronic Engineering from Tsinghua University in 2025.
My research focuses on physical agents and
long-horizon embodied tasks, with an emphasis on agent memory,
planning, and closed-loop control.
Email /
GitHub /
OpenReview
|
|
|
Tsinghua University
2025.09 - Present
- PhD student, Shenzhen International Graduate School
- Advisor: Prof. Yansong Tang
|
|
|
Tsinghua University
2021.09 - 2025.06
- B.Eng. in Electronic Engineering
- Department of Electronic Engineering
|
|
Tencent Robotics X
2026.05 - Present
- Research Intern
- Research on embodied agents for long-horizon tasks
|
Recent Publications
(* Equal contribution, † Corresponding author)
|
|
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
Shilin Ma, Chubin Zhang, Changyuan Wang, Yuji Wang, Yue Wu,
Zixuan Wang, Jingqi Tian, Zheng Zhu, Yansong Tang†
European Conference on Computer Vision (ECCV), 2026
[arXiv]
[Paper]
[Code]
[Project Page]
We propose a training-free, future-aware visual token pruning framework for efficient VLA manipulation.
|
|
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
Yixuan Zhu*, Shilin Ma*, Haolin Wang, Ao Li, Yanzhe Jing,
Yansong Tang†, Lei Chen, Jiwen Lu, Jie Zhou
International Conference on Learning Representations (ICLR), 2026
[arXiv]
[Paper]
[Code]
[Project Page]
We distill a visual autoregressive model into a one-step real-world image super-resolution model.
|
|
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
Yixuan Zhu*, Haolin Wang*, Shilin Ma*, Wenliang Zhao,
Yansong Tang†, Lei Chen, Jie Zhou†
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
[arXiv]
[Paper]
[Code]
We factorize diffusion features by frequency to balance structural consistency and appearance change in video editing.
|
|
Uncertainty-Aware World Model for Aerial Image-Goal Navigation
Deyi Zhu, Haoyu Fan, Yinan Zhu, Weichen Zhang, Shilin Ma,
Xinlei Chen, Yansong Tang
Preprint, 2026
[arXiv]
[Paper]
[Project Page]
We formulate trajectory scoring as conditional out-of-distribution detection for robust aerial navigation under future-state uncertainty.
|
|
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning
Jingqi Tian*, Haoji Zhang*, Lin Chen*, Hongbo Jin, Haonan Xu, Tianrui Zhu,
Xingming Shui, Shilin Ma, Wenjing Yang, Yansong Tang†
Preprint, 2026
[arXiv]
[Paper]
[Code]
[Project Page]
We propose an adaptive video-reasoning framework that learns when explicit reasoning is worth its generation cost.
|
|
MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models
Yue Wu, Changyuan Wang, Zixuan Wang, Shilin Ma, Yansong Tang
Preprint, 2026
[arXiv]
[Paper]
We develop a modality-aware post-training quantization framework that preserves cross-modal structure at 4-bit precision.
|
|