Label Science for Autonomous Vehicles
Autonomous vehicles and other physical AI systems depend on labels. A label is a set of spatial and contextual annotations applied to a vehicle's recorded driving footage. Precise annotations provide the ground-truth that helps a model understand the dynamic geometry of a scene.
Labels serve as helpful teaching guides that make AI smarter, more precise, and safer. Because labeling is a vital tool for teaching AI models about the real world, the industry has been working on ways to make the process more accurate and efficient.
Human labeling is expensive in hours.
It can take up to 800 human hours to label a single hour of driving footage.
AI-based auto-labeling is expensive in compute.
It can take up to $560 in compute to auto-label a single hour of driving footage.
When these expensive processes spend time labeling empty driving footage, the unhelpful footage that carries no label-relevant material, it is an unfortunate waste of time and money.
The result is that physical AI companies have decided to throw away their greatest asset. They are throwing away valuable footage generated by an expensive fleet of cars equipped with expensive sensors to capture every hour of driving footage. Having invested so much up-front capital to purchase an expensive fleet of autonomous vehicles, it would make sense to generate as much value as possible from it.
Our software analyzes your entire archive, ranks every scene, and moves the scenes with the highest ranking to the start of your queue.
Most recorded driving footage is ordinary and teaches nothing new. Physical AI is always searching for those rare scenes carrying learning opportunities related to safety and risk.
However, those safety-related scenes are rare and deeply embedded in thousands of hours of driving footage, which makes them hard to find.
WITNESS prioritizes those rare scenes by sorting your driving footage according to its Pattern Activity Index (PAI). This ranking gives you the most relevant scenes first, making your labeling pipeline more efficient.
You now have a ranked queue of highly relevant scenes instead of an unsorted pile of scenes.
With costs being so high to label a single hour of driving footage, that single hour of footage should be densely packed with label-relevant scenes.
WITNESS solves this problem by relying on neither AI nor humans to analyze, score, and rank a fleet’s recorded driving footage. With a much more cost-effective solution, WITNESS hands your labeling pipeline a ranked and ordered queue of scenes, densely packed with label-relevant material.
WITNESS uses its own proprietary software to analyze, score, rank, and sort your scenes in order of their “label-relevance,” which is to say their ability to help your model perceive and understand the physical world.
You send us your footage and we process it and send it back to you in a fully analyzed, scored, and ranked order of scenes that are sequenced such that:
WITNESS scores every scene, then moves the label-relevant scenes to the front of the labeling queue. The empty scenes are shuffled to the end of the line. No need to waste your expensive human hours or AI compute cycles on them.
As recorded
Densely Packed
Compared to the unsorted, unranked order, the first 5% of our PAI Ranking holds 8.4 times as many people. This category includes pedestrians and roadside workers.
The First Half
On average, when you have processed only 50% of our Ranked Order of Scenes, you will have reached 9 in 10 (90.3%) of every scene with a person in it.
The Second Half
Because we shuffle the label-relevant scenes to the front of the line, by the time you get to the second half of the archive, only a small amount of label-relevant material is in front of you: 1 in 10.
WITNESS is a software product that ingests multi-sensor fleet footage, ranks each scene in order of its instructional value to AI, and moves the label-relevant scenes to the front of the queue to be analyzed and labeled by AI models or by humans. An autonomous fleet now has a ranked queue of highly relevant scenes instead of an unsorted pile of scenes. WITNESS also produces actionable insights that optimize the labeling pipeline for maximum accuracy and efficiency.
The result: WITNESS finds the scenes with people in them 4.9 times faster than reviewing the same driving footage in the order it was recorded.
Pedestrians
The PAI Ranking prioritizes scenes that are relevant for labeling. One category that deserves special attention is pedestrians.
A Physical AI model must “get it right” when it comes to pedestrians, roadside workers, and anyone else on foot in the geometry of a scene.
Finding those scenes is vital to the safety of a driving model.
The six charts below measure the PAI Ranking on exactly that. They run on lidar alone. No detector, no model, no labels.
Accelerated discovery, measured
From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.
We cut the ranked archive into ten equal bands. Each band is 10,000 scenes. The first band is the 10,000 scenes we scored highest. The last band is the 10,000 we scored lowest. Then we counted, band by band, how many of those scenes have a person in them.
The dark bars do the same thing to the same archive, left in the order it came off the vehicle.
Read it left to right. Our bars start high and fall away to almost nothing. The recorded order stays flat, between 708 and 1,205 a band, because nothing sorted it and the people are spread evenly through it.
Inference-based labeling and human labeling both cost a great deal. WITNESS saves you real money and real time. What is it worth to get to market more quickly with a safer product?
| Ranking band, 10% each | PAI Ranking | Temporal Order | Multiple |
|---|---|---|---|
| Top band | 2,983 | 708 | 4.21x |
| 2nd band | 2,391 | 870 | 2.75x |
| 3rd band | 1,761 | 988 | 1.78x |
| 4th band | 1,323 | 1,104 | 1.20x |
| 5th band | 905 | 1,140 | 0.79x |
| 6th band | 557 | 1,205 | 0.46x |
| 7th band | 291 | 1,175 | 0.25x |
| 8th band | 122 | 1,027 | 0.12x |
| 9th band | 30 | 1,009 | 0.03x |
| Last band | 4 | 1,141 | 0.00x |
The chart below illustrates the power of accurate sorting. We move the label-relevant scenes to the front of the queue and we pack them tightly together. This saves you time. You spend your time processing label-relevant material rather than sifting through empty scenes that carry no instructional value for your model.
At the 10% mark our queue has handed you 2,983 scenes with people and the recorded order has handed you 708. Divide 2,983 by 708 and you get 4.21, the first point on the line. At the 20% mark our running totals are 5,374 against 1,578, which is 3.41. Carry on down the line.
Our ranking achieves a higher density of label-relevant scenes at the beginning of the queue. The empty scenes are moved to the back of the line. The point is that you start labeling at the front of the line and then you don’t need to waste your time labeling the archive that is populated mostly by empty scenes.
| How far into the queue | PAI Ranking | Temporal Order | Multiple |
|---|---|---|---|
| Top 10% | 2,983 | 708 | 4.21x |
| Top 20% | 5,374 | 1,578 | 3.41x |
| Top 30% | 7,135 | 2,566 | 2.78x |
| Top 40% | 8,458 | 3,670 | 2.30x |
| Top 50% | 9,363 | 4,810 | 1.95x |
| Top 60% | 9,920 | 6,015 | 1.65x |
| Top 70% | 10,211 | 7,190 | 1.42x |
| Top 80% | 10,333 | 8,217 | 1.26x |
| Top 90% | 10,363 | 9,226 | 1.12x |
| Top 100% | 10,367 | 10,367 | 1.00x |
This chart keeps a running total. Walk down the list and at every point it asks the same question: of all 10,367 scenes in the archive that have a person in them, how many have you met so far?
Our line climbs steeply and then flattens out, because most of the people are near the front. By the time you are 40% of the way down the list, 81.6 of every 100 scenes with a person are already behind you. Go 40% of the way through the footage in the order it was recorded and 35.4 of every 100 are.
You do not need an expensive model to get this. The ranking reads the vehicle's own lidar and nothing else.
Four bands in, you have reached most of the people in the archive. The rest of your footage is still there, in order, whenever you want it.
| Depth of the ranking | PAI Ranking | Temporal Order |
|---|---|---|
| Top 10% | 28.8% | 6.8% |
| Top 20% | 51.8% | 15.2% |
| Top 30% | 68.8% | 24.8% |
| Top 40% | 81.6% | 35.4% |
| Top 50% | 90.3% | 46.4% |
| Top 60% | 95.7% | 58.0% |
| Top 70% | 98.5% | 69.4% |
| Top 80% | 99.7% | 79.3% |
| Top 90% | 100.0% | 89.0% |
| Top 100% | 100.0% | 100.0% |
Say you have the time and the budget to review 5,000 scenes. That is one scene in every twenty.
Take the first 5,000 from our ranked queue and 1,525 of them have a person in them, which is 30 in every 100. Take the first 5,000 in the order the footage was recorded and 310 of them do, which is 6 in every 100. Across the whole archive the rate is 10 in every 100, so the recorded order actually starts below the archive's own average.
The same 5,000 scenes of work either way. Five times as many people in front of you at the end of it.
Accelerated discovery, measured
From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.
WITNESS is not inference based, it’s simply computational, which means we can process your footage and make it more densely packed with meaningful scenes before your model or your human labeling process wastes time and money.
| Order | Scenes reviewed | Holding a person | Share |
|---|---|---|---|
| PAI Ranking | 5,000 | 1,525 | 30.5% |
| Temporal Order | 5,000 | 310 | 6.2% |
| Whole archive, for reference | 100,000 | 10,367 | 10.4% |
Our ranking system makes your footage more densely packed with label-relevant scenes.
The chart before this one fixed the amount of work and counted what you got. This one turns the question around. It fixes what you want and counts the work.
Suppose your labeling pipeline needs 1,000 scenes with a person in them. Working down our queue you have them after opening 3,296 scenes. Working through the same footage in the order it was recorded you have them after opening 13,521. That is 10,225 fewer scenes opened for the same 1,000 found.
The number down the left is how many scenes with a person in them you want. The bars are how many scenes you have to open to get that many. Shorter bars are better.
Send us your entire archive or just enough to supply your labeling pipeline with the scenes it needs.
| Scenes with people wanted | PAI Ranking opens | Temporal Order opens | Ratio | Less footage opened |
|---|---|---|---|---|
| 250 | 1,135 | 3,810 | 3.36x | 70% |
| 500 | 1,887 | 7,477 | 3.96x | 75% |
| 1,000 | 3,296 | 13,521 | 4.10x | 76% |
| 2,000 | 6,515 | 24,279 | 3.73x | 73% |
| 3,000 | 10,058 | 33,991 | 3.38x | 70% |
| 5,000 | 18,244 | 51,639 | 2.83x | 65% |
Each band is 10,000 scenes, highest scored first. People are counted as scenes with a person in them. Objects are counted one by one.
Every chart so far has counted scenes with people. This one adds the other thing worth counting.
An object is one road user that somebody has to draw a box around: a car, a truck, a bus, a cyclist, a person. The archive holds 177,845 of them, spread across the 100,000 scenes.
The ranking is built to find people, so people is what it finds hardest. Objects follow along behind it, because a scene busy with people is usually busy with everything else as well. Both lines start far above an even spread and fall away together as you go down the list.
One pass serves both. The ranking is built to find people, and the objects come with them.
| Ranking band, 10% each | Scenes with people | Share of all | Admitted objects | Share of all |
|---|---|---|---|---|
| Top band | 2,983 | 28.8% | 33,305 | 18.7% |
| 2nd band | 2,391 | 23.1% | 29,546 | 16.6% |
| 3rd band | 1,761 | 17.0% | 25,175 | 14.2% |
| 4th band | 1,323 | 12.8% | 21,484 | 12.1% |
| 5th band | 905 | 8.7% | 18,204 | 10.2% |
| 6th band | 557 | 5.4% | 15,386 | 8.7% |
| 7th band | 291 | 2.8% | 12,993 | 7.3% |
| 8th band | 122 | 1.2% | 11,222 | 6.3% |
| 9th band | 30 | 0.3% | 9,411 | 5.3% |
| Last band | 4 | 0.0% | 1,119 | 0.6% |
Measured on the Zenseact Open Dataset: 100,000 annotated scenes, published by Zenseact and free for anyone to download and re-run. The ranking reads lidar only. It never opens a label file and never runs a model, so you can run it on footage you haven't labeled yet.
We did not build our own test set. A test set we assemble ourselves proves nothing about the archive sitting on a customer's servers. All three archives below are published by other organizations, carry labels those organizations placed, and can be downloaded by anyone who wants to run the same measurement we ran.
An autonomous vehicle archive of twenty-second driving clips carrying radar, lidar and camera. Its labels are placed by machine rather than by people, so it shows how WITNESS behaves when the answer sheet is itself automated.
A city driving archive carrying radar, lidar and camera. Its labels are placed by human annotators, which makes it the strictest test we run: our order is graded against people rather than against another machine.
The largest archive we have run, carrying camera and lidar. It has no radar, which is why our first pass on it is a lidar pass. It is released under a licence that permits commercial use, so a customer can download it and repeat our measurement without asking anyone's permission.
The result: reviewing one scene in twenty, a team working in PAI Ranking order reaches 4.9 times as many scenes with people in them as a team reviewing the same driving footage in the order it was recorded. The figure is measured on the Zenseact Open Dataset.
NVIDIA and nuScenes ask that numbers measured on their datasets not be published. So the figure we publish is measured on ZOD, which is released under a licence that permits it.
Zenseact Open Dataset (ZOD), © 2022 Zenseact AB, licensed under CC BY-SA 4.0.
We measure where you stand before we touch anything. What share of your existing labels are wrong, what the human work behind one accepted label costs you, and how much of the material worth labeling your current process finds.
We analyze your entire archive, rank every scene, and move the scenes with the highest ranking to the start of your queue. No query to type, no starter examples to provide; the rare, edge-case scenes surface first.
We reach the same scene-ranking quality with far fewer labels. WITNESS uses its PAI score, so ranking quality holds while the edge-case, label-relevant scenes are prioritized and presented for review ahead of the ordinary scenes.
We flag the automated, model-generated labels that are wrong. The industry assumes its labels are accurate, but published research shows otherwise. By flagging errors faster, WITNESS helps you troubleshoot your labeling pipeline.
A single mistake tells you little on its own. Read against your whole archive, that same mistake becomes one point in a larger pattern of where your model's understanding breaks down. WITNESS ranks that pattern and reveals the scenes your model handles worst, with no labels required.
We measure at the beginning and at the end of a single process. We report on what changed during our process. Both ends are measured the same way, so the improvement is a number your team can verify and validate.
The result accelerates discovery by 2.5x to 2.7x on all label-relevant material, and by 4.2x to 4.9x on the scenes with people in them. Faster through the same archive, and no more setting aside scenes that were label-relevant all along.