Xiyang Wu

Hi! I’m Xiyang Wu, a final-year Ph.D. student in Electrical and Computer Engineering at the University of Maryland, College Park, advised by Prof. Dinesh Manocha.

My research lies at the intersection of reinforcement learning, post-training, multimodal reasoning, and embodied AI. Broadly, I’m interested in building intelligent, reliable, and self-improving agents that can understand the world, learn from experience, and act effectively in complex environments.

More recently, I’ve been focusing on self-improving embodied agents and robotic systems. I’m exploring how multimodal reasoning, world models, and reinforcement learning can work together to help agents acquire reusable skills and become more autonomous and adaptive through physical interaction.

TL;DR I build self-improving AI agents that understand the physical world, learn reusable skills from experience, and act reliably in embodied environments.

prof_pic.png

University of Maryland, College Park

College Park, MD

Research Interests

My research revolves around three interconnected questions:

Understanding

How can agents understand physical environments, human intentions, and multimodal observations?

My work focuses on physical and social reasoning, multimodal grounding, world model evaluation, long-horizon video understanding, and synthetic data curation.

Learning

How can agents learn from experience and continuously improve through post-training and long-horizon interactions?

I study reinforcement learning, skill discovery, and self-improvement, with a focus on enabling agents to acquire reusable skills, learn from feedback, and solve increasingly complex tasks.

Key works

Acting

How can agents turn understanding and learning into robust decisions and actions in the physical world?

My work investigates robot planning, embodied decision-making, and VLA safety and robustness, with an emphasis on agents that adapt through physical interactions.

News

Sep 26, 2026 We release a technical report introducing ASSEMBLE, an atomic-skill framework for evidence-grounded long-video reasoning. With a 9B reader supervised by a 235B teacher, ASSEMBLE reaches 59.2 percent macro-averaged answer accuracy across three benchmarks, compared with 58.3 percent for Gemini-2.5-Pro, and improves overlap-based Grounded accuracy by 6.7 percent.
Aug 01, 2026 Two papers accepted by EMNLP 2026: MM-Zero enables vision-language models to self-evolve from zero human-annotated data through multi-role reinforcement learning; Graph-of-Skills introduces dependency-aware structural retrieval for scaling agents to massive skill libraries.
Jun 01, 2026 Two papers accepted by IROS 2026: SABER (Oral Presentation) red-teams VLA-controlled robots via stealthy instruction perturbations; FALCON introduces object-centric self-supervised pretraining for UAV action recognition.
Apr 07, 2026 We release a technical report introducing COS-PLAY, a co-evolution framework for long-horizon tasks where an LLM decision agent and a skill bank agent jointly improve through GRPO.
Mar 31, 2026 We release a technical report introducing SABER, an agentic black-box attack framework for red-teaming VLA-controlled robots.
Feb 01, 2026 Two papers accepted by CVPR 2026. First Frame Is the Place to Go for Video Content Customization proposes first-frame conditioning for efficient video customization; MASS introduces a physics-focused video benchmark and a model-agnostic method that injects 3D motion and spatiotemporal cues into VLMs.
Nov 23, 2025 We release a technical report introducing MASS-Bench, a physics-focused video benchmark, and MASS, a model-agnostic method that injects 3D motion and spatial-temporal cues into VLMs.
Sep 01, 2025 VideoHallu was accepted by NeurIPS 2025.
Aug 01, 2025 Advanced to Ph.D. candidate at the University of Maryland, College Park.
Jun 15, 2025 On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by IROS 2025.
May 02, 2025 We release VideoHallu, a benchmark for hallucinations in synthetic video understanding over common sense and physics.
Sep 01, 2024 AUTOHALLUSION was accepted by EMNLP 2024.
Jun 16, 2024 We release AUTOHALLUSION, an automatic benchmark generation approach for hallucination examples in vision-language models.
Jun 15, 2024 LANCAR and AGL-NET were accepted by IROS 2024.
Apr 01, 2024 On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by the VLADR Workshop at CVPR 2024.

Selected Publications

All publications
  1. arXiv 2026
    assemble.png
    ASSEMBLE: Atomic Skills for Evidence-Grounded Video Reasoning
    Xiyang Wu, Zongxia Li, Shengxin Zhang, Zhichao Liu, and Dinesh Manocha
    • Learning
    • RL Post-training
    • Video Reasoning
    • Skill Composition
    arXiv 2026
  2. arXiv 2026
    cosplay.png
    COS-PLAY: Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
    Xiyang Wu, Zongxia Li, Guangyao Shi, Alexander Duffy, Tyler Marques, Matthew Lyle Olson, Tianyi Zhou, and Dinesh Manocha
    • Learning
    • RL Post-training
    • Skill Learning
    • Long-horizon Agents
    arXiv 2026
  3. IROS 2026
    saber.png
    SABER: A Stealthy Agentic Black-Box Attack Framework for Vision-Language-Action Models
    Xiyang Wu, Guangyao Shi, Qingzi Wang, Zongxia Li, Amrit Singh Bedi, and Dinesh Manocha
    • Acting
    • VLA Models
    • Adversarial Robustness
    • Red Teaming
    IROS 2026 · Oral Presentation
  4. CVPR 2026
    mass.png
    MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
    Xiyang Wu*, Zongxia Li, Jihui Jin, Guangyao Shi, Gouthaman KV, Vishnu Raj, Nilotpal Sinha, Jingxi Chen, Fan Du, and Dinesh Manocha
    • Understanding
    • World Model
    • VLM
    • Video Understanding
    CVPR 2026
  5. IROS 2025
    adversary.png
    On the Vulnerability of LLM/VLM-Controlled Robotics
    Xiyang Wu, Souradip Chakraborty, Ruiqi Xian, Jing Liang, Tianrui Guan, Fuxiao Liu, Brian Sadler, Dinesh Manocha, and Amrit Singh Bedi
    • Acting
    • Embodied AI
    • Adversarial Robustness
    • AI Safety
    IROS 2025
  6. NeurIPS 2025
    videohallu.png
    VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
    Xiyang Wu*, Zongxia Li*, Yubin Qin, Guangyao Shi, Hongyang Du, Dinesh Manocha, Tianyi Zhou, and Jordan Lee Boyd-Graber
    • Understanding
    • World Model
    • VLM
    • Video Understanding
    NeurIPS 2025
  7. EMNLP 2024
    autohallusion.png
    AUTOHALLUSION: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
    Xiyang Wu*, Tianrui Guan*, Dianqi Li, Shuaiyi Huang, Xiaoyu Liu, Xijun Wang, Ruiqi Xian, Abhinav Shrivastava, Furong Huang, Jordan Lee Boyd-Graber, Tianyi Zhou, and Dinesh Manocha
    • Understanding
    • Synthetic Data Curation
    • Multimodal Evaluation
    • Hallucination
    EMNLP 2024
  8. IROS 2024
    lancar.png
    LANCAR: Leveraging Language for Context-Aware Robot Locomotion in Unstructured Environments
    Xiyang Wu*, Chak Lam Shek*, Wesley A. Suttle, Carl Busart, Erin Zaroukian, Dinesh Manocha, Pratap Tokekar, and Amrit Singh Bedi
    • Acting
    • Robot Learning
    • Reinforcement Learning
    • Language Grounding
    IROS 2024
  9. CVPR 2024
    hallusionbench.png
    HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
    Tianrui Guan*, Fuxiao Liu*, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou
    • Understanding
    • Multimodal Evaluation
    • Visual Reasoning
    • Hallucination
    CVPR 2024
  10. CoRL 2023
    iplan.png
    iPLAN: Intent-Aware Planning in Heterogeneous Traffic via Distributed Multi-Agent Reinforcement Learning
    Xiyang Wu, Rohan Chandra, Tianrui Guan, Amrit Singh Bedi, and Dinesh Manocha
    • Learning
    • Multi-agent RL
    • Intent-aware Planning
    • Autonomous Driving
    CoRL 2023 Oral (acceptance rate 6.6%). Best Paper and Presentation Award at the IROS 2023 workshop