We present EgoTAP, a heatmap-to-3D pose lifting method for highly accurate stereo egocentric 3D pose estimation. Severe self-occlusion and out-of-view limbs make accurate pose estimation challenging. Prior methods use joint heatmaps, but heatmap-to-3D conversion remains inaccurate.
We propose a Grid ViT Encoder that summarizes joint heatmaps into effective embeddings via self-attention, and a Propagation Network that uses skeletal structure to estimate obscure joints. EgoTAP reduces MPJPE by 23.9% vs. previous state of the art.