A sensor resolves less with range. The reference it is judged against is drawn by people reading that same degraded evidence. A model is characterised only under the conditions it was evaluated on, and the road supplies others. A system assembled from all four is given control of the vehicle.
For perception, simulation and planning teams · K-Radar, nuScenes, Waymo, NVIDIA
Four sources, none of them bounded
The sensor returns fewer points as range grows, fewer again in rain and occlusion. The reference is an instrument too: boxes a metre from the returns beneath them, one vehicle carrying four identities, a length inferred from a face no sensor saw. The model is characterised only where it was evaluated. The world supplies the rest.
These do not simply add. The annotator reads the same degraded returns the model does, so the reference is least certain exactly where the model is weakest. Measured against an unaudited reference, a residual understates the true error, and understates it most in the conditions that matter.
Test the human reference against the returns it was drawn from, correct what disagrees, and state what remains uncertain.
Measure how far each model departs from that reference, across five channels that behave differently from one another.
Fit those measurements into models of the error, conditioned on range, class, weather and occlusion.
Drive microscopic simulation with that fitted error, so agents act on perception that fails the way real perception fails.
Give motion planning the same model, so a decision can account for what perception does not promise.
What is measured
A single score ranks models against each other and says nothing about the situation a vehicle is in. These are distributions over range, class, weather and occlusion, so how much room does this planner need beside a truck at sixty metres in rain has an answer with a number in it. All five derive from one association matched on position alone: a class-aware match would refuse a pedestrian↔cyclist pair and conceal the misidentification it exists to measure.
Where the object is, against the enhanced reference.
How large it is. Length is the worst of the three: a lidar sees one end and infers the other.
Whether it is found at all, at this position.
Across a track, the fraction of frames in which it is found. A track is usually held throughout or lost early, so the mean hides two populations.
What it is called, on a pair matched by position alone.
The scene above is kradar_026 frames 60–99, played at the sensor's own 10 Hz. Each sweep is decimated from about 56,511 os2-64 returns to 14,000. The white boxes are the finalized reference version; the dashed boxes are a detector's output on the same frames. The diagrams here define the five channels and carry no fitted values: those are produced per recording inside the application.
Why it has to reach the decision
An uncertainty that is never quantified is not absorbed. It is inherited by the tracker, the predictor, and the decision.
A tracker binds a detection to an identity, a predictor extrapolates it, a planner commits to the result, each treating an estimate as a fact. A simulator fed exact detections validates a vehicle that will never exist. Both need the same object: the error as a distribution over the conditions that produced it.