Yutao Fan
Logo Harbin Institute of Technology
Logo Shanghai AI Laboratory

Hi! I am a first-year Ph.D. student in the School of Computer Science at Harbin Institute of Technology, supervised by Wangmeng Zuo and Lei Bai. I am also a joint research intern at the Shanghai Artificial Intelligence Laboratory, where I have the pleasure of working closely with Zhiheng Xi and Yiran Qin.

My research interests center on LLM agents, multimodal reasoning, and embodied manipulation, with a recent focus on reinforcement learning and foundation models for embodied agents.

Publications at NeurIPS 2025, ICLR 2025, ICML 2026, ACL 2026 (Oral), and TMLR.
Research spans LLM agents, reinforcement learning, multimodal reasoning, and embodied foundation models.
Building large-scale datasets and benchmarks for multimodal and scientific reasoning.

Research Focus
LLM Agents

Training and evaluating agentic systems that search, plan, collaborate, and adapt across dynamic environments.

Multimodal Reasoning

Improving how vision-language models understand ambiguous instructions, typography, and visual context.

Embodied Manipulation

Foundation models and world models that ground perception, language, and action for embodied agents.

Education
  • Harbin Institute of Technology & AI Lab
    Harbin Institute of Technology & AI Lab
    Ph.D. Student, advised by Prof. Wangmeng Zuo and Prof. Lei Bai
    Sep. 2025 - present
  • Harbin Institute of Technology
    Harbin Institute of Technology
    Bachelor in Software Engineering
    Sep. 2021 - Jul. 2025
Experience
  • Shanghai AI Laboratory
    Shanghai AI Laboratory
    Research Intern
    Oct. 2024 - present
Honors & Awards
  • Outstanding Student of Harbin Institute of Technology
    2025
  • Huawei Scholarship
    2024
News
2026
We released $N_0$-Foundation, $N_0$-VTLA, and $N_0$-TWAM, technical reports towards tactile intelligence for robotic manipulation!
Jul 26
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey is accepted by TMLR!
Jun 15
Our study on scaling behaviors of reinforcement learning for LLMs is accepted by ACL 2026 as an Oral!
May 15
SciAgentGym is accepted by ICML 2026!
May 01
Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study is accepted by ICML 2026!
May 01
2025
BMMR is accepted by NeurIPS 2025!
Sep 18
I started my Ph.D. at Harbin Institute of Technology, jointly working with Shanghai AI Laboratory.
Sep 01
Visual-O1 is accepted by ICLR 2025!
Jan 22
Selected Publications (view all )
$N_0$-Foundation: Towards the Age of Tactile Intelligence

NeoteAI Team, Fudan TEAI Team

arXiv coming soon. 2026 Tech Report Core Contributor

$N_0$-Foundation is a tactile-centric foundation for embodied manipulation that integrates tactile sensing hardware, large-scale multimodal data, hardware-agnostic tactile representation learning, and standardized evaluation into one system. Durable vision-based tactile sensors and $N_0$-TacUMI provide a synchronized acquisit...

$N_0$-Foundation: Towards the Age of Tactile Intelligence

NeoteAI Team, Fudan TEAI Team

arXiv coming soon. 2026 Tech Report Core Contributor

$N_0$-Foundation is a tactile-centric foundation for embodied manipulation that integrates tactile sensing hardware, large-scale multimodal data, hardware-agnostic tactile representation learning, and standardized evaluation into one system. Durable vision-based tactile sensors and $N_0$-TacUMI p...

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

NeoteAI Team, Fudan TEAI Team

2026 Tech Report Core Contributor

$N_0$-VTLA is a vision-tactile-language-action (VTLA) foundation model for fine-grained contact-rich manipulation. To our knowledge, it is the first VTLA model pretrained on tactile data at scale, learning broad contact priors from NeoData, our large-scale visuo-tactile robot dataset. A predictive tactile pathway distills the...

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

NeoteAI Team, Fudan TEAI Team

2026 Tech Report Core Contributor

$N_0$-VTLA is a vision-tactile-language-action (VTLA) foundation model for fine-grained contact-rich manipulation. To our knowledge, it is the first VTLA model pretrained on tactile data at scale, learning broad contact priors from NeoData, our large-scale visuo-tactile robot dataset. A predictiv...

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

NeoteAI Team, Fudan TEAI Team

2026 Tech Report Core Contributor

$N_0$-TWAM is a tactile-native world-action model for contact-rich manipulation that jointly predicts future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, pre-trained with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments an...

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation

NeoteAI Team, Fudan TEAI Team

2026 Tech Report Core Contributor

$N_0$-TWAM is a tactile-native world-action model for contact-rich manipulation that jointly predicts future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, pre-trained with visuo-tactile joint training over tactile-rich demonstratio...

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

Li Kang*, Yutao Fan*, Rui Li, Heng Zhou, Yiran Qin, Zhemeng Zhang, Songtao Huang, Xiufeng Song, Zaibin Zhang, Bruno N.Y. Chen, Zhenfei Yin, Dongzhan Zhou, Wangmeng Zuo, Lei Bai (* equal contribution)

preprint 2026 Preprint

CoEnv addresses coordination challenges in multi-agent robotic systems by introducing a compositional environment that integrates real-world and simulation components. The framework operates through three stages: digitizing physical workspaces, employing language models for action planning, and transferring validated strategi...

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

Li Kang*, Yutao Fan*, Rui Li, Heng Zhou, Yiran Qin, Zhemeng Zhang, Songtao Huang, Xiufeng Song, Zaibin Zhang, Bruno N.Y. Chen, Zhenfei Yin, Dongzhan Zhou, Wangmeng Zuo, Lei Bai (* equal contribution)

preprint 2026 Preprint

CoEnv addresses coordination challenges in multi-agent robotic systems by introducing a compositional environment that integrates real-world and simulation components. The framework operates through three stages: digitizing physical workspaces, employing language models for action planning, and t...

Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study
Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study

Zhiheng Xi*, Xin Guo*, Jiaqi Liu*, Jiazheng Zhang*, Yutao Fan, Zhihao Zhang, Shichun Liu, Mingxu Chai, Xiaowei Shi, Yitao Zhai, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang (* equal contribution)

International Conference on Machine Learning (ICML) 2026 Regular

We investigate whether reinforcement fine-tuning effectively enables LLM agents to generalize across scenarios. Our study examines three dimensions: within-environment generalization across task difficulty, cross-environment transfer to unseen environments, and sequential multi-environment training to assess both transfer and...

Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study
Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical Study

Zhiheng Xi*, Xin Guo*, Jiaqi Liu*, Jiazheng Zhang*, Yutao Fan, Zhihao Zhang, Shichun Liu, Mingxu Chai, Xiaowei Shi, Yitao Zhai, Xunliang Cai, Tao Gui, Qi Zhang, Xuanjing Huang (* equal contribution)

International Conference on Machine Learning (ICML) 2026 Regular

We investigate whether reinforcement fine-tuning effectively enables LLM agents to generalize across scenarios. Our study examines three dimensions: within-environment generalization across task difficulty, cross-environment transfer to unseen environments, and sequential multi-environment traini...

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

Minheng Ni, Yutao Fan, Zhengyuan Yang, Yeli Shen, Yuxiang Wei, Yaowen Zhang, Lijuan Wang, Lei Zhang, Wangmeng Zuo

preprint 2026 Preprint

CoEditor++ is a training-free framework for instruction-based image editing that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism. Built from publicly available components, CoEditor++ achieves competitive results on SmartEdit and AltBear compared ...

CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning

Minheng Ni, Yutao Fan, Zhengyuan Yang, Yeli Shen, Yuxiang Wei, Yaowen Zhang, Lijuan Wang, Lei Zhang, Wangmeng Zuo

preprint 2026 Preprint

CoEditor++ is a training-free framework for instruction-based image editing that decomposes editing into "what to edit" and "how to edit" through two cognitive stages with a reflective self-selection mechanism. Built from publicly available components, CoEditor++ achieves competitive results on S...

Better know nothing than half-know anything: A Precise and Efficient Dataset for Scientific Reasoning in Language Models
Better know nothing than half-know anything: A Precise and Efficient Dataset for Scientific Reasoning in Language Models

Yutao Fan*, Yizhou Wang*, Zhiheng Xi*, Lintao Wang, Jianyu Wu, Aoran Wang, Guanyu Li, Jiaqi Liu, Pengze Li, Heng Zhou, Jiayang Li, Wangmeng Zuo, Lei Bai, Shixiang Tang, Philip Torr, Zhenfei Yin (* equal contribution)

Under review. 2025

Large Language Models (LLMs) have achieved remarkable progress in reasoning tasks, i.e., coding and mathematics. However, their ability to perform scientific reasoning remains significantly limited, probably hampered by the scarcity of high-quality scientific reasoning datasets. Existing approaches either rely on LLM-generate...

Better know nothing than half-know anything: A Precise and Efficient Dataset for Scientific Reasoning in Language Models
Better know nothing than half-know anything: A Precise and Efficient Dataset for Scientific Reasoning in Language Models

Yutao Fan*, Yizhou Wang*, Zhiheng Xi*, Lintao Wang, Jianyu Wu, Aoran Wang, Guanyu Li, Jiaqi Liu, Pengze Li, Heng Zhou, Jiayang Li, Wangmeng Zuo, Lei Bai, Shixiang Tang, Philip Torr, Zhenfei Yin (* equal contribution)

Under review. 2025

Large Language Models (LLMs) have achieved remarkable progress in reasoning tasks, i.e., coding and mathematics. However, their ability to perform scientific reasoning remains significantly limited, probably hampered by the scarcity of high-quality scientific reasoning datasets. Existing approach...

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Zhiheng Xi*, Guanyu Li*, Yutao Fan*, Honglin Guo*, Yufang Liu, Xiaoran Fan, Jiaqi Liu, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai, Tao Ji, Tao Gui, Qi Zhang, Philip Torr, Xuanjing Huang (* equal contribution)

Annual Conference on Neural Information Processing Systems (NeurIPS) 2025 Poster

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple-choice, fill-in-the-blank, an...

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Zhiheng Xi*, Guanyu Li*, Yutao Fan*, Honglin Guo*, Yufang Liu, Xiaoran Fan, Jiaqi Liu, Jingchao Ding, Wangmeng Zuo, Zhenfei Yin, Lei Bai, Tao Ji, Tao Gui, Qi Zhang, Philip Torr, Xuanjing Huang (* equal contribution)

Annual Conference on Neural Information Processing Systems (NeurIPS) 2025 Poster

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 110k college-level questions spanning 300 UNESCO-defined subjects, spanning diverse formats-multiple...

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Minheng Ni, Yutao Fan, Lei Zhang, Wangmeng Zuo

The International Conference on Learning Representations (ICLR) 2025 Poster

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. However, even highly intelligent l...

Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Minheng Ni, Yutao Fan, Lei Zhang, Wangmeng Zuo

The International Conference on Learning Representations (ICLR) 2025 Poster

As large-scale models evolve, language instructions are increasingly utilized in multi-modal tasks. Due to human language habits, these instructions often contain ambiguities in real-world scenarios, necessitating the integration of visual context or common sense for accurate interpretation. Howe...

All publications