Compare Google's system from five years ago.[1] (at 7:42) This new video is at n00b level compared to that. [1] https://embed.ted.com/talks/chris_urmson_how_a_driverless_ca...
The camera is bird's eye on the google presentation. And it is a recording. So jitter/flicker can be cleaned up/smoothed out and the data can be massaged in ways that a real time system may not be able to do. This is a presentation. I'd be /very/ surprised if this animation was from RAW data as-is. Also there seems to be Lidar data (point clouds) which Tesla doesn't have. So while this means bounding boxes may have l…
I’m not sure what is implied with saying it’s a recording - both the Google and Tesla presentations are “recordings” and equal opportunity to pick best case examples, but I would bet strongly there is nothing not RAW = “real time“ for their respective compute platforms.
The top down viewpoint helps show off the quality (still by no means perfect) of the world representation. If you projected Tesla’s model into 3D you would see far more jitter than in the video overlay for a variety of reasons.
That said, I think comparing them directly on specific technical components is a bit of a sidebar. They are taking two very different paths along the way to a still ambiguous problem. Both are leading their respective approaches, but have fundamentally different and unproven assumptions.
Also worth looking not just at how accurately objects are detected but what the visualizations show about the intent of other road users. The Google video shows predicted trajectories for important objects in a number of scenes. We don’t get to see any of that clearly from Tesla, and that is by no means a small part of the problem. Not sure if it is there and not shown, just highlighting there is a lot more downstream even once you are finding objects reliably in the sensors.