Lifeng Zhuo | 卓立峰

I’m an undergraduate in Software Engineering at the School of Computer Science, Shanghai Jiao Tong University, advised by Prof. Chuan Wen and Prof. Cewu Lu. My research focuses on robot learning and multimodal policies for contact-rich manipulation.

Lifeng Zhuo's avatar

Publications

FA-RDP preserves diverse approach trajectories before contact and switches to rapid force feedback after contact.

FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation

Lifeng Zhuo*, Wendi Chen*, Han Xue, Shirun Tang, Jun Lv, Cewu Lu, Chuan Wen

arXiv preprint, 2026 · * Equal contribution

FA-RDP addresses the trade-off between diverse motion planning and rapid force feedback in contact-rich manipulation. A shared visual-force Transformer supports both low-frequency, multi-step diffusion sampling before contact and high-frequency, single-step control after contact. A learned multimodality indicator selects the appropriate frequency, while Manifold Consistency Distillation enables fast action prediction on the robot action manifold. Experiments on three manipulation tasks demonstrate improved success rates while preserving diverse pre-contact trajectories.

IPR-1 learns shared physical and causal mechanisms from over 1,000 games and transfers to unseen games.

IPR-1: Interactive Physical Reasoner

Mingyu Zhang*, Lifeng Zhuo*, Tianxi Tan, Guocan Xie, Xian Nie, Yan Li, Renjie Zhao, Zizhu He, Ziyu Wang, Jiting Cai, Yong-Lu Li

CVPR 2026 · * Equal contribution

IPR-1 learns physical and causal reasoning through interaction with over 1,000 heterogeneous games. It combines a vision-language policy with world-model rollouts that anticipate action outcomes and guide reinforcement learning. A physics-centric action representation, PhysCode, connects semantic intent with physical dynamics. The model is evaluated from basic physical intuition to goal-directed reasoning, improves with more training games and interaction steps, and transfers to unseen games without additional training.

RESBev predicts clean BEV features from temporal context and fuses them with corrupted features to restore perception.

RESBev: Making BEV Perception More Robust

Lifeng Zhuo, Kefan Jin, Zhe Liu, Hesheng Wang

NeurIPS 2026

RESBev improves the robustness of bird's-eye-view perception by treating feature recovery as a latent semantic prediction problem. A temporal world model uses historical BEV observations and ego-motion to predict clean feature priors, which are fused with corrupted observations to reconstruct the scene representation. The module integrates into existing perception systems without changing their backbone. Experiments on nuScenes demonstrate improved resilience to natural sensor disturbances and adversarial attacks with limited fine-tuning.