An overview of our model structure. The input point cloud will first voxelized and go through three of ours 3D backbone layer to produce RoIs. Then, the convolution features and the predicted foreground score will be used in further refinement stage. In refinement stage, we will utilize the score and calculate the distance to the corresponding grid points as the sample criteria.