Haozhe Jia 贾浩哲

Incoming Ph.D. Student at Peking University EPIC Lab · Research Intern at Galbot

I am an incoming Ph.D. student at Peking University. I have joined EPIC Lab, where I work under the supervision of Prof. He Wang and Prof. Li Yi. I am currently a research intern at Galbot and an undergraduate at Shandong University.

My research focuses on humanoid robot learning and control, particularly on connecting generative motion models with physical execution. I am interested in how robots can translate language instructions into coordinated, physically grounded whole-body behaviors. My work spans motion generation, dynamics-aware representations, and language-conditioned humanoid control.

Portrait of Haozhe Jia

Research

Humanoid Learning & Control

Translating language instructions into coordinated whole-body behaviors through dynamics-aware representations and robot control.

Generative Motion Modeling

Learning language-conditioned motion representations and generative models for expressive and controllable movement.

Physics-Grounded Generative Models

Incorporating physical structure into generative modeling to improve consistency and generalization.

Selected Publications

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

Under ReviewFirst author

Introduces dynamics-aligned joint intent representations that anticipate support transfer, contact switching, and balance preparation for streaming language-conditioned humanoid control.

ECHO: Edge-Cloud Humanoid Orchestration for Language-to-Motion Control

Under ReviewFirst author

Connects cloud-based motion generation with on-device closed-loop tracking for language-to-motion humanoid control, validated in simulation and on real hardware.

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

ICML 2026First author

Proposes REPA-P to align denoising features with physics-aware representations, improving physical consistency and out-of-distribution robustness.

ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model

ACM MM 2025Co-first author

Introduces step-aware temporal modulation and late-stage CFG reduction for diffusion motion models, improving semantic alignment and retrieval performance.

RMDM: Physics-Informed Representation Alignment for Sparse Radio-Map Reconstruction

ACM MM 2025 OralFirst author

Combines a PINN-based field initializer with a diffusion refiner to reconstruct sparse radio maps accurately under physically constrained settings.

More Publications

DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space

ICML 2025Second co-author

Develops an end-to-end diffusion model in DCT space for higher-resolution image generation with stronger efficiency and spectral interpretability.

LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model

Under review at ECCVFirst author

Uses temporal semantic anchors and low-frequency motion anchors to improve deep U-Net alignment, gradient flow, and convergence in diffusion motion synthesis.

Towards Better Evaluation Metrics for Text-to-Motion Generation

WWW 2026Co-first author

Introduces OTMS and MMMD, two evaluation metrics designed to better correlate text-to-motion quality with human judgment.

POLARIS: Projection-Orthogonal Least Squares for Robust and Adaptive Inversion in Diffusion Models

arXiv preprintCollaborating author

Derives a theoretically grounded projection schedule for diffusion inversion, improving reconstruction quality without substantial extra computation.

RadioFlow: Efficient Radio Map Construction Framework with Flow Matching

Under review at TCNNFirst author

Uses deterministic flow matching for fast radio map construction, reducing parameter count and inference time while maintaining reconstruction quality.

Free-T2M: Frequency Enhanced Text-to-Motion Diffusion Model With Consistency Loss

Under review at ICRAFirst author

Introduces frequency-aware consistency supervision to stabilize motion denoising and improve semantic fidelity in diffusion-based text-to-motion generation.

Guided Path Sampling: Steering Diffusion Models Back on Track with Principled Path Guidance

WWW 2026Collaborating author

Applies manifold-aware interpolation and dynamic guidance schedules to keep diffusion sampling closer to valid data trajectories.

Experience

Current

Research Intern

Galbot

2025.12 - 2026.06 Beijing, China

Embodied AI Algorithm Intern

LimX Dynamics

  • Developed ECHO, a language-driven humanoid motion control system with a compact 38-DoF action representation.
  • Built a cloud-edge streaming pipeline: cloud diffusion generates motion references; on-device lightweight controller performs closed-loop tracking.
  • Validated in MuJoCo simulation and on real humanoid hardware.
2024.12 - 2025.12 Guangzhou, China

Research Assistant

Hong Kong University of Science and Technology (Guangzhou)

  • Led research on PhyRMDM, Free-T2M, and LUMA, spanning radio map reconstruction and text-driven human motion generation.
  • Owned the full pipeline from model selection and training to ablation design; all code open-sourced.
  • First-author / co-first-author publications at ICML and ACM MM (Oral).

Contact

Leave a Note

Feel free to leave a message or discuss collaboration ideas.

If comments do not load below, read or leave a note on GitHub.