Hi! I’m Xiyang Wu, a final-year Ph.D. student in Electrical and Computer Engineering at the University of Maryland, College Park, advised by Prof. Dinesh Manocha.
My research lies at the intersection of reinforcement learning, post-training, multimodal reasoning, and embodied AI. Broadly, I’m interested in building intelligent, reliable, and self-improving agents that can understand the world, learn from experience, and act effectively in complex environments.
More recently, I’ve been focusing on self-improving embodied agents and robotic systems. I’m exploring how multimodal reasoning, world models, and reinforcement learning can work together to help agents acquire reusable skills and become more autonomous and adaptive through physical interaction.
TL;DR I build self-improving AI agents that understand the physical world, learn reusable skills from experience, and act reliably in embodied environments.
University of Maryland, College Park
College Park, MD
Research Interests
My research revolves around three interconnected questions:
Understanding
How can agents understand physical environments, human intentions, and multimodal observations?
My work focuses on physical and social reasoning, multimodal grounding, world model evaluation, long-horizon video understanding, and synthetic data curation.
Key works
Learning
How can agents learn from experience and continuously improve through post-training and long-horizon interactions?
I study reinforcement learning, skill discovery, and self-improvement, with a focus on enabling agents to acquire reusable skills, learn from feedback, and solve increasingly complex tasks.
Acting
How can agents turn understanding and learning into robust decisions and actions in the physical world?
My work investigates robot planning, embodied decision-making, and VLA safety and robustness, with an emphasis on agents that adapt through physical interactions.
Key works
News
| Sep 26, 2026 | We release a technical report introducing ASSEMBLE, an atomic-skill framework for evidence-grounded long-video reasoning. With a 9B reader supervised by a 235B teacher, ASSEMBLE reaches 59.2 percent macro-averaged answer accuracy across three benchmarks, compared with 58.3 percent for Gemini-2.5-Pro, and improves overlap-based Grounded accuracy by 6.7 percent. |
|---|---|
| Aug 01, 2026 | Two papers accepted by EMNLP 2026: MM-Zero enables vision-language models to self-evolve from zero human-annotated data through multi-role reinforcement learning; Graph-of-Skills introduces dependency-aware structural retrieval for scaling agents to massive skill libraries. |
| Jun 01, 2026 | Two papers accepted by IROS 2026: SABER (Oral Presentation) red-teams VLA-controlled robots via stealthy instruction perturbations; FALCON introduces object-centric self-supervised pretraining for UAV action recognition. |
| Apr 07, 2026 | We release a technical report introducing COS-PLAY, a co-evolution framework for long-horizon tasks where an LLM decision agent and a skill bank agent jointly improve through GRPO. |
| Mar 31, 2026 | We release a technical report introducing SABER, an agentic black-box attack framework for red-teaming VLA-controlled robots. |
| Feb 01, 2026 | Two papers accepted by CVPR 2026. First Frame Is the Place to Go for Video Content Customization proposes first-frame conditioning for efficient video customization; MASS introduces a physics-focused video benchmark and a model-agnostic method that injects 3D motion and spatiotemporal cues into VLMs. |
| Nov 23, 2025 | We release a technical report introducing MASS-Bench, a physics-focused video benchmark, and MASS, a model-agnostic method that injects 3D motion and spatial-temporal cues into VLMs. |
| Sep 01, 2025 | VideoHallu was accepted by NeurIPS 2025. |
| Aug 01, 2025 | Advanced to Ph.D. candidate at the University of Maryland, College Park. |
| Jun 15, 2025 | On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by IROS 2025. |
| May 02, 2025 | We release VideoHallu, a benchmark for hallucinations in synthetic video understanding over common sense and physics. |
| Sep 01, 2024 | AUTOHALLUSION was accepted by EMNLP 2024. |
| Jun 16, 2024 | We release AUTOHALLUSION, an automatic benchmark generation approach for hallucination examples in vision-language models. |
| Jun 15, 2024 | LANCAR and AGL-NET were accepted by IROS 2024. |
| Apr 01, 2024 | On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by the VLADR Workshop at CVPR 2024. |