arXiv Computer Vision#tech
VideoMDM: Towards 3D Human Motion Generation From 2D Supervisiontranslating…
Factuality: 64/100UnknownarXiv Publishing Co.
arXiv:2606.13364v1 Announce Type: cross
Abstract: We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild as