News

Sep 26, 2026 We release a technical report introducing ASSEMBLE, an atomic-skill framework for evidence-grounded long-video reasoning. With a 9B reader supervised by a 235B teacher, ASSEMBLE reaches 59.2 percent macro-averaged answer accuracy across three benchmarks, compared with 58.3 percent for Gemini-2.5-Pro, and improves overlap-based Grounded accuracy by 6.7 percent.
Aug 01, 2026 Two papers accepted by EMNLP 2026: MM-Zero enables vision-language models to self-evolve from zero human-annotated data through multi-role reinforcement learning; Graph-of-Skills introduces dependency-aware structural retrieval for scaling agents to massive skill libraries.
Jun 01, 2026 Two papers accepted by IROS 2026: SABER (Oral Presentation) red-teams VLA-controlled robots via stealthy instruction perturbations; FALCON introduces object-centric self-supervised pretraining for UAV action recognition.
Apr 07, 2026 We release a technical report introducing COS-PLAY, a co-evolution framework for long-horizon tasks where an LLM decision agent and a skill bank agent jointly improve through GRPO.
Mar 31, 2026 We release a technical report introducing SABER, an agentic black-box attack framework for red-teaming VLA-controlled robots.
Feb 01, 2026 Two papers accepted by CVPR 2026. First Frame Is the Place to Go for Video Content Customization proposes first-frame conditioning for efficient video customization; MASS introduces a physics-focused video benchmark and a model-agnostic method that injects 3D motion and spatiotemporal cues into VLMs.
Nov 23, 2025 We release a technical report introducing MASS-Bench, a physics-focused video benchmark, and MASS, a model-agnostic method that injects 3D motion and spatial-temporal cues into VLMs.
Sep 01, 2025 VideoHallu was accepted by NeurIPS 2025.
Aug 01, 2025 Advanced to Ph.D. candidate at the University of Maryland, College Park.
Jun 15, 2025 On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by IROS 2025.
May 02, 2025 We release VideoHallu, a benchmark for hallucinations in synthetic video understanding over common sense and physics.
Sep 01, 2024 AUTOHALLUSION was accepted by EMNLP 2024.
Jun 16, 2024 We release AUTOHALLUSION, an automatic benchmark generation approach for hallucination examples in vision-language models.
Jun 15, 2024 LANCAR and AGL-NET were accepted by IROS 2024.
Apr 01, 2024 On the Vulnerability of LLM/VLM-Controlled Robotics was accepted by the VLADR Workshop at CVPR 2024.
Feb 26, 2024 HallusionBench was accepted by CVPR 2024. Data and code are on GitHub.
Feb 15, 2024 We release a technical report on the robustness and safety of integrating LLMs and VLMs into robotics. Project page: adversary-vlm-robotics.
Oct 23, 2023 We release an early report and analysis on failure modes of GPT-4V and LLaVA-1.5, later released as HallusionBench.
Oct 01, 2023 iPLAN received the Best Paper Award from the MRS Workshop at IROS 2023.
Aug 01, 2023 iPLAN was accepted by CoRL 2023 as an Oral (acceptance rate 6.6%).
Jul 01, 2023 One paper was accepted by Digital Signal Processing.
Aug 01, 2021 Started Ph.D. at the University of Maryland, College Park.
Aug 01, 2019 Started M.S. at the Georgia Institute of Technology.