Portrait
Songheng Yin
Robotics System Engineer
Mujin, Inc
About Me

I am a robotics engineer at Mujin@Tokyo. I have been working broadly on computer vision, artificial intelligence, and robotics. I am now also looking for PhD opportunities starting in 2027 or 2028 chiikawa

During my undergrad, I was a research student at Pair Lab@Georgia Tech, advised by Prof. Animesh Garg. I am very fortunate to have worked with him and Dr. Wei Yu.

On the industry side, I have interned at Amazon@Seattle and SMBC@NYC, and worked at Tencent@Beijing as a backend infrastructure engineer.

I have lived in Nanjing, Toronto, Vancouver, New York, Beijing, and now Tokyo. In my free time, I enjoy landscape photography and learning languages. I always love making new friends and finding ways to work together — feel free to reach out~ chiikawa

Education
  • Columbia University
    Columbia University
    M.Sc. in Computer Science
    2022 - 2024
  • University of Toronto
    University of Toronto
    B.Sc. in Computer Science
    2018 - 2022
Honors & Awards
  • Norman Stuart Robertson Scholarship (Top 3 in math major)
  • University of Toronto Research Excellence Award (UTEA)
  • College Silver Medal

Publications

Conferences

EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning

Wei Yu, Songheng Yin, Steve Easterbrook, Animesh Garg

International Conference on Learning Representations (ICLR) 2025

EgoSim generates egocentric exploration videos of virtual worlds conditioned on multi-modal inputs.

EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning

Wei Yu, Songheng Yin, Steve Easterbrook, Animesh Garg

International Conference on Learning Representations (ICLR) 2025

EgoSim generates egocentric exploration videos of virtual worlds conditioned on multi-modal inputs.

Modular Action Concept Grounding in Semantic Video Prediction

Wei Yu, Wenxin Chen, Songheng Yin, Steve Easterbrook, Animesh Garg

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022

We propose modular action concept grounding (MAC), which represents actions as compositions of grounded semantic concepts to improve semantic video prediction.

Modular Action Concept Grounding in Semantic Video Prediction

Wei Yu, Wenxin Chen, Songheng Yin, Steve Easterbrook, Animesh Garg

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022

We propose modular action concept grounding (MAC), which represents actions as compositions of grounded semantic concepts to improve semantic video prediction.

Preprint / Under Review

Spatial Transport of Integration Error in Generative ODEs

Songheng Yin, Wei Yu, Souleyman Boudouh

arXiv preprint 2026

Few-step generative sampling error is not merely local—it is shaped by how signed truncation errors are transported across regions by the model’s learned dynamics.

Spatial Transport of Integration Error in Generative ODEs

Songheng Yin, Wei Yu, Souleyman Boudouh

arXiv preprint 2026

Few-step generative sampling error is not merely local—it is shaped by how signed truncation errors are transported across regions by the model’s learned dynamics.

QUOTA: Quantization via Output-channel Targeted Allocation

Souleyman Boudouh, Simla Burcu Harma, Alireza Khodamoradi, Abdulrahman Mahmoud, Songheng Yin, Kristof Denolf, Babak Falsafi

under review

Learned, per-channel mixed-precision budget can largely amortize the collapse across diverse backbones and scales.

QUOTA: Quantization via Output-channel Targeted Allocation

Souleyman Boudouh, Simla Burcu Harma, Alireza Khodamoradi, Abdulrahman Mahmoud, Songheng Yin, Kristof Denolf, Babak Falsafi

under review

Learned, per-channel mixed-precision budget can largely amortize the collapse across diverse backbones and scales.

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models

Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg

arXiv preprint 2026

MosaicMem introduces a hybrid spatial memory that enables controllable video world models.

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models

Wei Yu, Runjia Qian, Yumeng Li, Liquan Wang, Songheng Yin, Sri Siddarth Chakaravarthy P, Dennis Anthony, Yang Ye, Yidi Li, Weiwei Wan, Animesh Garg

arXiv preprint 2026

MosaicMem introduces a hybrid spatial memory that enables controllable video world models.