Ablation studies of the backbone structure adoption on the mod level of the KITTI val set and all the results are evaluated with average precision and calculated by 40 recall positions. Also, we provide FPS for the comparison of the model efficiency.
| Methods | Car | 3D Moderate Pedestrian | Cyclist | FPS |
|---|---|---|---|---|
| Voxel-RCNN, [6] | 84.87 | 59.87 | 73.23 | 30 |
| FocalConv*, [4] | 85.26 | 60.04 | 73.08 | 19 |
| Ours | 85.34 | 60.30 | 74.33 | 29 |
| Methods | Car | 3D Moderate Pedestrian | Cyclist | FPS |
|---|---|---|---|---|
| Voxel-RCNN, [ | 84.87 | 59.87 | 73.23 | 30 |
| FocalConv | 85.26 | 60.04 | 73.08 | 19 |
| Ours | 85.34 | 60.30 | 74.33 | 29 |
Note: The above results are reproduced by the publicly release model ([28]).
:Note that we only adopt the attention backbone structure while the remaining part is the same as the baseline model.
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.