Humanoid Learning & Control
Translating language instructions into coordinated whole-body behaviors through dynamics-aware representations and robot control.
Incoming Ph.D. Student at Peking University EPIC Lab · Research Intern at Galbot
I am an incoming Ph.D. student at Peking University. I have joined EPIC Lab, where I work under the supervision of Prof. He Wang and Prof. Li Yi. I am currently a research intern at Galbot and an undergraduate at Shandong University.
My research focuses on humanoid robot learning and control, particularly on connecting generative motion models with physical execution. I am interested in how robots can translate language instructions into coordinated, physically grounded whole-body behaviors. My work spans motion generation, dynamics-aware representations, and language-conditioned humanoid control.
Translating language instructions into coordinated whole-body behaviors through dynamics-aware representations and robot control.
Learning language-conditioned motion representations and generative models for expressive and controllable movement.
Incorporating physical structure into generative modeling to improve consistency and generalization.
Introduces dynamics-aligned joint intent representations that anticipate support transfer, contact switching, and balance preparation for streaming language-conditioned humanoid control.
Connects cloud-based motion generation with on-device closed-loop tracking for language-to-motion humanoid control, validated in simulation and on real hardware.
Proposes REPA-P to align denoising features with physics-aware representations, improving physical consistency and out-of-distribution robustness.
Introduces step-aware temporal modulation and late-stage CFG reduction for diffusion motion models, improving semantic alignment and retrieval performance.
Combines a PINN-based field initializer with a diffusion refiner to reconstruct sparse radio maps accurately under physically constrained settings.
Develops an end-to-end diffusion model in DCT space for higher-resolution image generation with stronger efficiency and spectral interpretability.
Uses temporal semantic anchors and low-frequency motion anchors to improve deep U-Net alignment, gradient flow, and convergence in diffusion motion synthesis.
Introduces OTMS and MMMD, two evaluation metrics designed to better correlate text-to-motion quality with human judgment.
Derives a theoretically grounded projection schedule for diffusion inversion, improving reconstruction quality without substantial extra computation.
Uses deterministic flow matching for fast radio map construction, reducing parameter count and inference time while maintaining reconstruction quality.
Introduces frequency-aware consistency supervision to stabilize motion denoising and improve semantic fidelity in diffusion-based text-to-motion generation.
Applies manifold-aware interpolation and dynamic guidance schedules to keep diffusion sampling closer to valid data trajectories.
Galbot
LimX Dynamics
Hong Kong University of Science and Technology (Guangzhou)
Feel free to leave a message or discuss collaboration ideas.
If comments do not load below, read or leave a note on GitHub.