Ego3DPose: Capturing 3D Cues from Binocular Egocentric Views

Binocular 3D pose from stereo correspondence and perspective cues.

Taeho Kang1, Kyungjin Lee1, Jinrui Zhang2, Youngki Lee1

1Seoul National University · 2Central South University

Presented at SIGGRAPH Asia 2023

Abstract

We present Ego3DPose, a highly accurate binocular egocentric 3D pose reconstruction system. The binocular egocentric setup offers practicality and usefulness in various applications, however, it remains largely under-explored. It has been suffering from low pose estimation accuracy due to viewing distortion, severe self-occlusion, and limited field-of-view of the joints in egocentric 2D images. Here, we notice that two important 3D cues, stereo correspondences, and perspective, contained in the egocentric binocular input are neglected. Current methods heavily rely on 2D image features, implicitly learning 3D information, which introduces biases towards commonly observed motions and leads to low overall accuracy. We observe that they not only fail in challenging occlusion cases but also in estimating visible joint positions.

To address these challenges, we propose two novel approaches. First, we design a two-path network architecture with a path that estimates pose per limb independently with its binocular heatmaps. Without full-body information provided, it alleviates bias toward trained full-body distribution. Second, we use the egocentric view of body limbs, which exhibits strong perspective variance (e.g., a significantly large-size hand when it is close to the camera). We propose a new perspective-aware representation using trigonometry, enabling the network to estimate the 3D orientation of limbs. Finally, we develop an end-to-end pose reconstruction network that synergizes both techniques. Our comprehensive evaluations demonstrate that Ego3DPose outperforms state-of-the-art models by a pose estimation error (i.e., MPJPE) reduction of 23.1% against the UnrealEgo baseline in the UnrealEgo dataset. Our qualitative results highlight the superiority of our approach across a range of scenarios and challenges.

Problem

Two fisheye cameras on a pair of glasses see the body from a poor angle, and prior methods miss even the limbs those cameras show clearly.

Egocentric 3D pose reconstruction lets avatar creation and motion analysis leave the room, because the cameras travel with the wearer. Glasses with two fisheye cameras are the most practical rig, but the field of view is restricted and the body occludes itself, so accuracy sits far below third-person estimation.

Estimated 3D pose reprojected onto the binocular input, from UnrealEgo and from Ego3DPose

Estimated 3D pose reprojected onto the input pair. Top, Akada et al. 2022 (UnrealEgo); bottom, Ego3DPose.

The green skeleton in each circle is the estimated 3D pose reprojected onto the left and right inputs. In the top row UnrealEgo places the circled arm away from where the image shows it, although its own heatmaps found the joint in 2D; the bottom row is our estimate on the same frames. The failure is most common where body parts move independently and a full-body prior cannot place one limb.

Key idea

Binocular egocentric images carry two 3D cues that prior networks leave unused, stereo correspondence and perspective, and each has to enter the network in a particular form.

Stereo correspondence is the first. Triangulating estimated 2D joints is fragile, because egocentric occlusion and distortion push 2D error straight into 3D. So stereo has to be learned alongside 2D estimation rather than run after it, and learned per limb, from that limb's own features, so it does not lean on the training set's full-body pose distribution. That is the Stereo Matcher: a weight-shared network that reads one limb's binocular features and outputs that limb's 3D orientation, which assists the lifting instead of replacing it.

Perspective is the second. A fisheye view from the head shrinks the lower limbs and enlarges shoulders and upper arms. Prior work treats this as distortion to train through, but the same foreshortening measures how far a limb points along the viewing direction: in the figure above, an upper arm much wider at the top than at the bottom points nearly perpendicular to the camera's view plane. Ego3DPose writes that orientation into each limb's 2D heatmap with trigonometry, so the heatmap keeps its uncertainty.

Method

Ego3DPose keeps the standard encoder–decoder pose network and adds two modules: a Perspective Embedding Heatmap at feature extraction and a Stereo Matcher at the encoder stage.

Perspective cue

View Plane Wide Narrow

The limb-view angle θl.

The perspective effect is stored as the limb-view angle θl between a limb and the camera's view plane. In the drawing on the left, the upper dashed line is the view plane through the camera and the lower one its parallel through the shoulder; θl is how far the child joint rises out of that plane, seen from its parent. On the right, the same arm in a real fisheye frame is wide at the shoulder and narrow toward the wrist.

Perspective Embedding Heatmap

Perspective Embedding Heatmap Cos Channel Sin Channel c cos θl c sin θl θl c

Two heatmap channels hold the angle as a vector.

One 2D heatmap has to hold where the limb is, how confident the estimate is, and the limb-view angle, without one value overwriting another. Each Perspective Embedding Heatmap therefore carries two channels, the scaled sine and cosine of θl. On the left is one real heatmap with a single pixel marked in red; read as a vector, on the right, that pixel's direction is θl and its length c is the confidence, so perspective stays entangled with probability.

Stereo Matcher and pose decoder

Perspective Embedding Heatmaps Left View Right View Stereo Matcher Shared Weights θ 3D Orientations Limb i fθ oi Limb j fθ oj ⋮ ⋮ ⋮ ⋮ 14 Limbs Joint Position Heatmaps Heatmap Encoder Pose Features Pose Decoder

Stereo and pose features travel separately and meet at the decoder.

Two U-Nets produce the Perspective Embedding Heatmaps per limb and the Joint Position Heatmaps per joint, in both views. The Stereo Matcher takes one limb's heatmaps at a time and predicts that limb's 3D orientation oi, the unit vector from parent to child; it need not know which limb it is, so the weights are shared across all 14 limbs. The heatmap encoder compresses all the heatmaps into pose features, which meet the orientations at the pose decoder. The orientations are detached from the decoder's gradient, so the Stereo Matcher learns from stereo alone.

Results

On UnrealEgo, Ego3DPose lowers MPJPE by 23.1% against UnrealEgo and 27.0% against EgoGlass, and the two modules gain more together than either does alone.

In the clips, red is the ground truth and grey is the predicted pose. The binocular input is at the far left, then EgoGlass (Zhao et al. 2021), UnrealEgo (Akada et al. 2022) and Ego3DPose.

UnrealEgo

EgoCap

UnrealEgo

Synthetic UnrealEgo, all motion categories; MPJPE and PA-MPJPE in millimetres. In the MPJPE column, 60.82 against UnrealEgo's 79.08 is the 23.1% reduction.

Method MPJPE ↓ PA-MPJPE ↓
EgoGlass83.3361.56
UnrealEgo79.0859.26
Ego3DPose (Ours)60.8248.47

EgoCap

Real-world EgoCap, on its 3D-annotated portion. The margins are smaller: EgoCap has fewer severe self-occlusions and a narrower range of poses.

Method MPJPE ↓ PA-MPJPE ↓
EgoGlass61.7846.06
UnrealEgo59.1648.22
Ego3DPose (Ours)54.4140.24

Ablation

UnrealEgo dataset. B is the UnrealEgo baseline network, PH adds the Perspective Embedding Heatmap and SM adds the Stereo Matcher; in B+SM the Stereo Matcher reads the two adjacent joint heatmaps instead.

Method MPJPE ↓ PA-MPJPE ↓
B (UnrealEgo)79.0859.26
B+PH75.8258.52
B+SM66.7252.29
B+PH+SM (Ours)60.8248.47

Against B in the MPJPE column, the heatmap alone lowers error by 4.1%, the Stereo Matcher alone by 15.6%, both together by 23.1%: the perspective heatmap pays off mainly as input to the Stereo Matcher.

Limitations

Lower-body occlusion, camera-specific distortion and a narrow real-world test set remain open.

Occlusion remains the hard case, since many egocentric motions hide the lower body; temporal optimisation and inverse kinematics are the next steps. A Stereo Matcher that reads stereo correspondence can come to depend on the rig, so different optics may show larger errors. And the real-world evaluation uses only the 3D-annotated portion of EgoCap: few subjects and frames, standing activities, a green-screen lab.

BibTeX

@inproceedings{10.1145/3610548.3618147,
  author    = {Kang, Taeho and Lee, Kyungjin and Zhang, Jinrui and Lee, Youngki},
  title     = {Ego3DPose: Capturing 3D Cues from Binocular Egocentric Views},
  year      = {2023},
  month     = {December},
  isbn      = {9798400703157},
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  url       = {https://doi.org/10.1145/3610548.3618147},
  doi       = {10.1145/3610548.3618147},
  booktitle = {SIGGRAPH Asia 2023 Conference Papers},
  pages     = {1--10},
  articleno = {82},
  series    = {SA '23}
}