Pose estimation and point-cloud perception (undergraduate thesis + Jilin University lab)

Contents

Pose estimation and point-cloud perception (undergraduate thesis + Jilin University lab)

Pose estimation

My undergraduate thesis focused on object pose estimation. Taking an RGB image and the 3D model of the detected object as input, I mapped 2D image pixels onto the 3D point cloud of the model’s surface. From that correspondence I regressed the object’s pose using PnP (Perspective-n-Point) and RANSAC. The approach also pairs with a deep-learning refinement step that improves accuracy further once PnP has produced an initial pose. The experiments showed that establishing an explicit mapping between the 2D plane and 3D space yields more accurate pose estimates than the comparable methods.

Overall algorithm framework
Overall algorithm framework

Generating a synthetic dataset from randomly posed 3D models on COCO backgrounds
Generating a synthetic dataset from randomly posed 3D models on COCO backgrounds

Dataset generation results
Dataset generation results

The principle of UV mapping
The principle of UV mapping

The mapping between the UV map and the object's surface point cloud
The mapping between the UV map and the object’s surface point cloud

Design of the UV map generation network
Design of the UV map generation network

UV map generation results
UV map generation results

UV map generation results compared against the calibrated images
UV map generation results compared against the calibrated images

Generated point cloud compared against the calibrated point cloud
Generated point cloud compared against the calibrated point cloud

Overall approach: initial pose regression with RANSAC + PnP
Overall approach: initial pose regression with RANSAC + PnP

Design of the deep-learning pose regression network
Design of the deep-learning pose regression network

Pose recognition results
Pose recognition results

Point-cloud perception

In my senior year I worked mainly on point-cloud perception in the lab at Jilin University, starting with converting data from the Livox format into KITTI format.

Data format conversion
Data format conversion

I then trained PointPillars on the Livox dataset and implemented forward inference.

Forward inference and recognition, result 1
Forward inference and recognition, result 1
Forward inference and recognition, result 2
Forward inference and recognition, result 2

The experience gave me a systematic view of point-cloud processing methods, and a much clearer sense of where purely point-cloud approaches fall short in real object recognition. That is what got me interested in fusing point-cloud and visual perception.

Purely point-cloud methods still have obvious room for improvement in model recognition
Purely point-cloud methods still have obvious room for improvement in model recognition