Harbin Institute of Technology
Shanghai AI LaboratoryHi! I am a first-year Ph.D. student in the School of Computer Science at Harbin Institute of Technology, supervised by Wangmeng Zuo and Lei Bai. I am also a joint research intern at the Shanghai Artificial Intelligence Laboratory, where I have the pleasure of working closely with Zhiheng Xi and Yiran Qin.
My research interests center on LLM agents, multimodal reasoning, and embodied manipulation, with a recent focus on reinforcement learning and foundation models for embodied agents.
Training and evaluating agentic systems that search, plan, collaborate, and adapt across dynamic environments.
Improving how vision-language models understand ambiguous instructions, typography, and visual context.
Foundation models and world models that ground perception, language, and action for embodied agents.
$N_0$-Foundation is a tactile-centric foundation for embodied manipulation that integrates tactile sensing hardware, large-scale multimodal data, hardware-agnostic tactile representation learning, and standardized evaluation into one system. Durable vision-based tactile sensors and $N_0$-TacUMI provide a synchronized acquisit...
$N_0$-Foundation is a tactile-centric foundation for embodied manipulation that integrates tactile sensing hardware, large-scale multimodal data, hardware-agnostic tactile representation learning, and standardized evaluation into one system. Durable vision-based tactile sensors and $N_0$-TacUMI p...
$N_0$-VTLA is a vision-tactile-language-action (VTLA) foundation model for fine-grained contact-rich manipulation. To our knowledge, it is the first VTLA model pretrained on tactile data at scale, learning broad contact priors from NeoData, our large-scale visuo-tactile robot dataset. A predictive tactile pathway distills the...
$N_0$-VTLA is a vision-tactile-language-action (VTLA) foundation model for fine-grained contact-rich manipulation. To our knowledge, it is the first VTLA model pretrained on tactile data at scale, learning broad contact priors from NeoData, our large-scale visuo-tactile robot dataset. A predictiv...
$N_0$-TWAM is a tactile-native world-action model for contact-rich manipulation that jointly predicts future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, pre-trained with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments an...
$N_0$-TWAM is a tactile-native world-action model for contact-rich manipulation that jointly predicts future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, pre-trained with visuo-tactile joint training over tactile-rich demonstratio...

CoEnv addresses coordination challenges in multi-agent robotic systems by introducing a compositional environment that integrates real-world and simulation components. The framework operates through three stages: digitizing physical workspaces, employing language models for action planning, and transferring validated strategi...
CoEnv addresses coordination challenges in multi-agent robotic systems by introducing a compositional environment that integrates real-world and simulation components. The framework operates through three stages: digitizing physical workspaces, employing language models for action planning, and t...

We investigate whether reinforcement fine-tuning effectively enables LLM agents to generalize across scenarios. Our study examines three dimensions: within-environment generalization across task difficulty, cross-environment transfer to unseen environments, and sequential multi-environment training to assess both transfer and...
We investigate whether reinforcement fine-tuning effectively enables LLM agents to generalize across scenarios. Our study examines three dimensions: within-environment generalization across task difficulty, cross-environment transfer to unseen environments, and sequential multi-environment traini...

CoEditor++ is a training-free framework for instruction-based image editing that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism. Built from publicly available components, CoEditor++ achieves competitive results on SmartEdit and AltBear compared ...
CoEditor++ is a training-free framework for instruction-based image editing that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism. Built from publicly available components, CoEditor++ achieves competitive results on S...

Large Language Models (LLMs) have achieved remarkable progress in reasoning tasks, i.e., coding and mathematics. However, their ability to perform scientific reasoning remains significantly limited, probably hampered by the scarcity of high-quality scientific reasoning datasets. Existing approaches either rely on LLM-generate...
Large Language Models (LLMs) have achieved remarkable progress in reasoning tasks, i.e., coding and mathematics. However, their ability to perform scientific reasoning remains significantly limited, probably hampered by the scarcity of high-quality scientific reasoning datasets. Existing approach...

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple-choice, fill-in-the-blank, an...
In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple...

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent l...
As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. Howe...