Hang Yin | 尹航

Hang Yin is currently a PhD student in the Department of Automation, Tsinghua University, advised by Prof. Jie Zhou and Prof. Jiwen Lu. In 2023, he received his Bachelor's degree in the Department of Automation, Tsinghua University.

I work on embodied AI and robot learning. My research focuses on:

  • Vision-language-action models that scale a single policy across diverse tasks, environments and robot embodiments, covering both navigation and manipulation.
  • Reinforcement learning for embodied agents that studies real-world RL and lifelong self-learning, where an agent keeps improving from its own interaction through persistent multimodal memory and without human supervision.
  • My research also covers:
  • Training-free navigation that leverages 3D scene graphs and LLM reasoning for object-goal, goal-oriented and vision-and-language navigation.
  • Email  /  Google Scholar  /  Github

    profile photo
    News

  • 2026-06: One paper on lifelong navigation is released on arXiv.
  • 2025-08: One paper on vision-and-language navigation is accepted to CoRL 2025.
  • 2025-02: One paper on goal-oriented navigation is accepted to CVPR 2025.
  • 2024-09: One paper on object-goal navigation is accepted to NeurIPS 2024.
  • Publications
    dise AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning
    Hang Yin*, Yinan Liang*, Jiazhao Zhang*, Jiahang Liu, Minghan Li, Zhizheng Zhang, He Wang
    arXiv, 2026
    [arXiv] [Project Page]

    We propose AllDayNav, a lifelong navigation system via real-world reinforcement learning. The robot autonomously builds a self-evolving multimodal memory, generates self-instructions, and continuously improves its navigation policy without human supervision, achieving near-100% success rates.

    dise GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
    Hang Yin*, Haoyu Wei*, Xiuwei Xu, Wenxuan Guo, Jie Zhou, Jiwen Lu
    Conference on Robot Learning (CoRL), 2025
    [arXiv] [Code] [Project Page]

    We propose a training-free framework for vision-and-language navigation. Our framework formulates navigation guidance as graph constraint optimization by decomposing instructions into explicit spatial constraints, enabling zero-shot adaptation to unseen environments.

    dise UniGoal: Towards Universal Zero-shot Goal-oriented Navigation
    Hang Yin*, Xiuwei Xu*, Linqing Zhao, Ziwei Wang, Jie Zhou, Jiwen Lu
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
    [arXiv] [Code] [Project Page] [中文解读]

    We propose UniGoal, a unified graph representation for zero-shot goal-oriented navigation. Based on online 3D scene graph prompting for LLM, our method can be directly applied to different kinds of scenes and goals without training.

    dise SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
    Hang Yin*, Xiuwei Xu*, Zhenyu Wu, Jie Zhou, Jiwen Lu
    Thirty-eighth Conference on Neural Information Processing Systems (NeurIPS), 2024
    [arXiv] [Code] [Project Page] [中文解读]

    We propose a training-free object-goal navigation framework by leveraging LLM and VFMs. We construct an online hierarchical 3D scene graph and prompt LLM to exploit structure information contained in subgraphs for zero-shot decision making.

    dise F2F-AP: Flow-to-Future Asynchronous Policy for Real-time Dynamic Manipulation
    Haoyu Wei, Xiuwei Xu, Ziyang Cheng, Hang Yin, Angyuan Ma, Bingyao Yu, Jie Zhou, Jiwen Lu
    arXiv, 2026
    [arXiv] [Code] [Project Page]

    We propose F2F-AP, a flow-to-future asynchronous policy for real-time dynamic manipulation. It predicts object flow to synthesize future observations and aligns visual features with future states, allowing policies to compensate for latency and interact with moving objects.

    dise AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation
    Wenxuan Guo, Xiuwei Xu, Yichen Liu, Xiangyu Li, Hang Yin, Huangxing Chen, Wenzhao Zheng, Jianjiang Feng, Jie Zhou, Jiwen Lu
    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
    [arXiv] [Project Page]

    We propose AwareVLN, a self-aware reasoning framework for vision-language navigation. It triggers structured reasoning at key navigation nodes to understand scene context, task progress, and next-step plans, improving instruction following in simulation and real-world navigation.

    dise MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
    Zhenyu Wu, Angyuan Ma, Xiuwei Xu, Hang Yin, Yinan Liang, Ziwei Wang, Jiwen Lu, Haibin Yan
    Conference on Robot Learning (CoRL), 2025
    [arXiv] [Project Page]

    We propose a general framework for mobile manipulation, which can be divided into docking point selection and fixed-base manipulation. We model the docking point selection stage as an optimization process, to let the agent move and touch target keypoint under several constraints.

    dise IGL-Nav: Incremental 3D Gaussian Localization for Image-goal Navigation
    Wenxuan Guo, Xiuwei Xu, Hang Yin, Ziwei Wang, Jianjiang Feng, Jie Zhou, Jiwen Lu
    International Conference on Computer Vision (ICCV), 2025
    [arXiv] [Code] [Project Page]

    We propose IGL-Nav, an incremental 3D Gaussian localization framework for image-goal navigation. It supports challenging scenarios where the camera for goal capturing and the agent's camera have very different intrinsics and poses, e.g., a cellphone and a RGB-D camera.


    Website Template


    © Hang Yin | Last update: Sep. 5, 2026