Yu Qi Northeastern University

Ph.D. Student in Computer Science

I am a rising fourth-year CS Ph.D. student at Northeastern University, advised by Prof. Lawson L. S. Wong. Before that, I received my B.S. in Computer Science from Peking University in the year of 2023.

Yu Qi

Research

My research focuses on the post-training — SFT and RL — of robot policy models. More specifically, I am interested in the following questions:

News

Selected Publications

Full list on Google Scholar.

CoRL
2026
Yu Qi, Z. Ye, X. Xu, Y. Lu, A. Sandhu, B. Hu, H. Huang, J. Tremblay, L. L. S. Wong
Conference on Robot Learning (CoRL), 2026
A diagnostic framework and post-training data-collection recipe for compositional generalization, attributing policy shortcuts to individual instruction factors such as color, verb, object, and spatial relation.
arXiv
2026
H. Huang, Z. Ye, L. Zhao, B. Hu, M. Jia, Yu Qi, A. Agha, D. Wang, R. Platt, et al.
arXiv preprint, 2026
A 3D closed-loop manipulation policy that predicts actions as pixel classification over camera-plane action maps, avoiding the token explosion of discretized action spaces.
CVPRW
2026
Z. Qian, X. Chi, Yu Qi, H. Li, Z.-Y. Chen, S. Zhang
Conference on Computer Vision and Pattern Recognition (CVPR), Workshop, 2026
A reinforcement learning framework that trains the world model and action model jointly through online interaction, pushing world-action policies beyond their demonstration data.
ECCV
2026
Yu Qi, X. Xu, Z. Guo, S. Ma, R. Zhang, X. Chen, R. An, R. Xing, J. Zhang, H. Huang, P.-A. Heng, J. Tremblay, L. L. S. Wong
European Conference on Computer Vision (ECCV), 2026
A video reasoning benchmark for reasoning coherence — whether generated events stay causally consistent across frames — evaluated with and without text and visual hints.
ICLR
2026
X. Zhu, Yu Qi, Y. Zhu, R. Walters, R. Platt
International Conference on Learning Representations (ICLR), 2026
An SE(3)-equivariant multi-task transformer for language-conditioned 3D manipulation, guaranteeing that policy behavior stays consistent under rotations and translations of the scene.
CVPR
2026
Z. Guo*, X. Chen*, R. Zhang*, R. An*, Yu Qi*, D. Jiang, X. Li, M. Zhang, H. Li, et al.
Conference on Computer Vision and Pattern Recognition (CVPR), Findings, 2026
A benchmark and empirical study of video generation models as zero-shot reasoners, spanning spatial, geometric, physical, temporal, and embodied reasoning.
ICML
2026
Yu Qi*, H. Zhao*, Z. Guo*, S. Ma, Z. Chen, Y. Han, R. Zhang, Z. Lin, S. Xin, et al.
International Conference on Machine Learning (ICML), 2026
A skill-level embodied benchmark, plus BEAR-Agent, an embodied agent that calls visual and spatial reasoning tools — drawing on its own observations — to solve embodied tasks.
RA-L
2025
Y. Zhu, Z. Ye, B. Hu, H. Zhao, Yu Qi, D. Wang, R. Platt
IEEE Robotics and Automation Letters (RA-L), 2025
A visuotactile policy that predicts residual in-hand rotation with an SO(2)-equivariant network over surface normals reconstructed from tactile images.
RA-L
2025
H. Zhao, Yu Qi, B. Hu, Y. Zhu, Z. Chen, H. Tian, X. Zhu, O. Howell, H. Huang, et al.
IEEE Robotics and Automation Letters (RA-L), 2025
A hierarchical skill-learning framework that uses object-centric skills to connect a high-level vision-language model with low-level visuomotor policies.
ICML
2025
D. Jiang, R. Zhang, Z. Guo, Y. Li, Yu Qi, X. Chen, L. Wang, J. Jin, C. Guo, S. Yan, et al.
International Conference on Machine Learning (ICML), 2025
A benchmark for chain-of-thought reasoning in large multimodal models, measuring reasoning quality, robustness, and efficiency across six domains.
CVPR
2025
Yu Qi*, Y. Ju*, T. Wei, C. Chu, L. L. S. Wong, H. Xu
Conference on Computer Vision and Pattern Recognition (CVPR), 2025
A large-scale dataset and multi-task policy for daily pairwise object assembly, covering everyday tasks such as plugging into sockets and arranging flowers in vases.
CoRL
2024
Y. Qian, X. Zhu, O. Biza, S. Jiang, L. Zhao, H. Huang, Yu Qi, R. Platt
Conference on Robot Learning (CoRL), 2024
A vision-language grasping system that reasons with GPT-4o about which object to clear next, uncovering and grasping targets in heavy clutter.
arXiv
2026
H. Huang, L. Zhao, H. Liu, Z. Ye, S.-Y. Huang, M. Jia, B. Hu, F. Lin, Yu Qi, et al.
arXiv preprint, 2026
An imitation-learning method that represents manipulation as image-space keypoint trajectories, recovering 3D end-effector poses by triangulation and enabling equivariant augmentation.

Academic Services