A human eye in close-up, gazing toward a field of stars, with the galaxy reflected in the iris.
The Technology

WITNESS

Label Science for Autonomous Vehicles

8x ACCELERATED DISCOVERY OF LABEL-RELEVANT SCENES
The Top 5% of our ranked scenes holds 8.4 times as many people as the first 5% of the unranked, unsorted footage.
Label Science

Autonomous vehicles and other physical AI systems depend on labels. A label is a set of spatial and contextual annotations applied to a vehicle's recorded driving footage. Precise annotations provide the ground-truth that helps a model understand the dynamic geometry of a scene.

Labels serve as helpful teaching guides that make AI smarter, more precise, and safer. Because labeling is a vital tool for teaching AI models about the real world, the industry has been working on ways to make the process more accurate and efficient.

The Problem
01 Human Hours

Human labeling is expensive in hours.

It can take up to 800 human hours to label a single hour of driving footage.

02 AI Compute Costs

AI-based auto-labeling is expensive in compute.

It can take up to $560 in compute to auto-label a single hour of driving footage.

When these expensive processes spend time labeling empty driving footage, the unhelpful footage that carries no label-relevant material, it is an unfortunate waste of time and money.

The result is that physical AI companies have decided to throw away their greatest asset. They are throwing away valuable footage generated by an expensive fleet of cars equipped with expensive sensors to capture every hour of driving footage. Having invested so much up-front capital to purchase an expensive fleet of autonomous vehicles, it would make sense to generate as much value as possible from it.

The Solution

WITNESS Software

Our software analyzes your entire archive, ranks every scene, and moves the scenes with the highest ranking to the start of your queue.

Most recorded driving footage is ordinary and teaches nothing new. Physical AI is always searching for those rare scenes carrying learning opportunities related to safety and risk.

However, those safety-related scenes are rare and deeply embedded in thousands of hours of driving footage, which makes them hard to find.

WITNESS prioritizes those rare scenes by sorting your driving footage according to its Pattern Activity Index (PAI). This ranking gives you the most relevant scenes first, making your labeling pipeline more efficient.

You now have a ranked queue of highly relevant scenes instead of an unsorted pile of scenes.

Efficient Ranking

Model-Free | Human-Free

With costs being so high to label a single hour of driving footage, that single hour of footage should be densely packed with label-relevant scenes.

WITNESS solves this problem by relying on neither AI nor humans to analyze, score, and rank a fleet’s recorded driving footage. With a much more cost-effective solution, WITNESS hands your labeling pipeline a ranked and ordered queue of scenes, densely packed with label-relevant material.

WITNESS uses its own proprietary software to analyze, score, rank, and sort your scenes in order of their “label-relevance,” which is to say their ability to help your model perceive and understand the physical world.

How It Works

You send us your footage and we process it and send it back to you in a fully analyzed, scored, and ranked order of scenes that are sequenced such that:

  1. The most instructional, label-relevant scenes land at the top of your list, ready for your review. They help your model perceive and understand the world.
  2. The empty scenes are shuffled to the bottom of the list so you don’t waste time on them. Your AI models and your human labeling pipelines spend their time efficiently labeling the scenes that matter most.
Ranked and Sorted

An unsorted archive becomes a ranked queue.

WITNESS scores every scene, then moves the label-relevant scenes to the front of the labeling queue. The empty scenes are shuffled to the end of the line. No need to waste your expensive human hours or AI compute cycles on them.

Label-relevant scenes Empty scenes

As recorded

Label the meaningful scenes first.
The empty scenes won’t waste time clogging your expensive labeling pipeline.

Schematic. The bar counts illustrate our methodology and are not a measure of exact distribution. Measured results vary from fleet to fleet and from one condition to the next.

8.4x

Densely Packed

Compared to the unsorted, unranked order, the first 5% of our PAI Ranking holds 8.4 times as many people. This category includes pedestrians and roadside workers.

9 in 10

The First Half

On average, when you have processed only 50% of our Ranked Order of Scenes, you will have reached 9 in 10 (90.3%) of every scene with a person in it.

1 in 10

The Second Half

Because we shuffle the label-relevant scenes to the front of the line, by the time you get to the second half of the archive, only a small amount of label-relevant material is in front of you: 1 in 10.

WITNESS is a software product that ingests multi-sensor fleet footage, ranks each scene in order of its instructional value to AI, and moves the label-relevant scenes to the front of the queue to be analyzed and labeled by AI models or by humans. An autonomous fleet now has a ranked queue of highly relevant scenes instead of an unsorted pile of scenes. WITNESS also produces actionable insights that optimize the labeling pipeline for maximum accuracy and efficiency.

The result: WITNESS finds the scenes with people in them 4.9 times faster than reviewing the same driving footage in the order it was recorded.

Three figures walking across a field of light

Pedestrians

The PAI Ranking prioritizes scenes that are relevant for labeling. One category that deserves special attention is pedestrians.

A Physical AI model must “get it right” when it comes to pedestrians, roadside workers, and anyone else on foot in the geometry of a scene.

Finding those scenes is vital to the safety of a driving model.

The six charts below measure the PAI Ranking on exactly that. They run on lidar alone. No detector, no model, no labels.

5x

Accelerated discovery, measured

From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.

WITNESS RANKING

Bringing Order to Chaos

We cut the ranked archive into ten equal bands. Each band is 10,000 scenes. The first band is the 10,000 scenes we scored highest. The last band is the 10,000 we scored lowest. Then we counted, band by band, how many of those scenes have a person in them.

The dark bars do the same thing to the same archive, left in the order it came off the vehicle.

Read it left to right. Our bars start high and fall away to almost nothing. The recorded order stays flat, between 708 and 1,205 a band, because nothing sorted it and the people are spread evenly through it.

PAI RankingTemporal Order, the order it came off the vehicle
08001,6002,4003,200PAI Ranking, Top tenth: 2,983 scenes with peopleTemporal Order, Top tenth: 708 scenes with peoplePAI Ranking, 2nd tenth: 2,391 scenes with peopleTemporal Order, 2nd tenth: 870 scenes with peoplePAI Ranking, 3rd tenth: 1,761 scenes with peopleTemporal Order, 3rd tenth: 988 scenes with peoplePAI Ranking, 4th tenth: 1,323 scenes with peopleTemporal Order, 4th tenth: 1,104 scenes with peoplePAI Ranking, 5th tenth: 905 scenes with peopleTemporal Order, 5th tenth: 1,140 scenes with peoplePAI Ranking, 6th tenth: 557 scenes with peopleTemporal Order, 6th tenth: 1,205 scenes with peoplePAI Ranking, 7th tenth: 291 scenes with peopleTemporal Order, 7th tenth: 1,175 scenes with peoplePAI Ranking, 8th tenth: 122 scenes with peopleTemporal Order, 8th tenth: 1,027 scenes with peoplePAI Ranking, 9th tenth: 30 scenes with peopleTemporal Order, 9th tenth: 1,009 scenes with peoplePAI Ranking, Last tenth: 4 scenes with peopleTemporal Order, Last tenth: 1,141 scenes with peopleTop2nd3rd4th5th6th7th8th9thLastTenth of the ranking, best first
Each band is 10,000 scenes, one tenth of the archive. The counts are scenes with a person in them.
The takeaway

Inference-based labeling and human labeling both cost a great deal. WITNESS saves you real money and real time. What is it worth to get to market more quickly with a safer product?

Table view
Ranking band, 10% eachPAI RankingTemporal OrderMultiple
Top band2,9837084.21x
2nd band2,3918702.75x
3rd band1,7619881.78x
4th band1,3231,1041.20x
5th band9051,1400.79x
6th band5571,2050.46x
7th band2911,1750.25x
8th band1221,0270.12x
9th band301,0090.03x
Last band41,1410.00x
4.2x

Front-Loaded Scene Selection

The chart below illustrates the power of accurate sorting. We move the label-relevant scenes to the front of the queue and we pack them tightly together. This saves you time. You spend your time processing label-relevant material rather than sifting through empty scenes that carry no instructional value for your model.

At the 10% mark our queue has handed you 2,983 scenes with people and the recorded order has handed you 708. Divide 2,983 by 708 and you get 4.21, the first point on the line. At the 20% mark our running totals are 5,374 against 1,578, which is 3.41. Carry on down the line.

Scenes with a person, PAI Ranking divided by Temporal Order1.0x, the same in both orders
1x2x3x4x1.0x, the same archive in both ordersTop 10% of the queue: 4.21 times as many scenes with a person as the same depth of Temporal Order4.21x10%Top 20% of the queue: 3.41 times as many scenes with a person as the same depth of Temporal Order3.41x20%Top 30% of the queue: 2.78 times as many scenes with a person as the same depth of Temporal Order30%Top 40% of the queue: 2.30 times as many scenes with a person as the same depth of Temporal Order2.30x40%Top 50% of the queue: 1.95 times as many scenes with a person as the same depth of Temporal Order50%Top 60% of the queue: 1.65 times as many scenes with a person as the same depth of Temporal Order1.65x60%Top 70% of the queue: 1.42 times as many scenes with a person as the same depth of Temporal Order70%Top 80% of the queue: 1.26 times as many scenes with a person as the same depth of Temporal Order80%Top 90% of the queue: 1.12 times as many scenes with a person as the same depth of Temporal Order90%Top 100% of the queue: 1.00 times as many scenes with a person as the same depth of Temporal Order1.00x100%How far into the queue you have gone
Counted on the 10,367 scenes in the archive that hold a person. Higher is better, and 1.0 means the two orders are level.
The takeaway

Our ranking achieves a higher density of label-relevant scenes at the beginning of the queue. The empty scenes are moved to the back of the line. The point is that you start labeling at the front of the line and then you don’t need to waste your time labeling the archive that is populated mostly by empty scenes.

Table view
How far into the queuePAI RankingTemporal OrderMultiple
Top 10%2,9837084.21x
Top 20%5,3741,5783.41x
Top 30%7,1352,5662.78x
Top 40%8,4583,6702.30x
Top 50%9,3634,8101.95x
Top 60%9,9206,0151.65x
Top 70%10,2117,1901.42x
Top 80%10,3338,2171.26x
Top 90%10,3639,2261.12x
Top 100%10,36710,3671.00x
82%

Where the people sit in the ranking

This chart keeps a running total. Walk down the list and at every point it asks the same question: of all 10,367 scenes in the archive that have a person in them, how many have you met so far?

Our line climbs steeply and then flattens out, because most of the people are near the front. By the time you are 40% of the way down the list, 81.6 of every 100 scenes with a person are already behind you. Go 40% of the way through the footage in the order it was recorded and 35.4 of every 100 are.

You do not need an expensive model to get this. The ranking reads the vehicle's own lidar and nothing else.

PAI RankingTemporal Order, the order it came off the vehicle
0%25%50%75%100%Temporal Order: the top 10% holds 6.8% of every scene with a personTemporal Order: the top 20% holds 15.2% of every scene with a personTemporal Order: the top 30% holds 24.8% of every scene with a personTemporal Order: the top 40% holds 35.4% of every scene with a personTemporal Order: the top 50% holds 46.4% of every scene with a personTemporal Order: the top 60% holds 58.0% of every scene with a personTemporal Order: the top 70% holds 69.4% of every scene with a personTemporal Order: the top 80% holds 79.3% of every scene with a personTemporal Order: the top 90% holds 89.0% of every scene with a personTemporal Order: the top 100% holds 100.0% of every scene with a personPAI Ranking: the top 10% holds 28.8% of every scene with a personPAI Ranking: the top 20% holds 51.8% of every scene with a personPAI Ranking: the top 30% holds 68.8% of every scene with a personPAI Ranking: the top 40% holds 81.6% of every scene with a personPAI Ranking: the top 50% holds 90.3% of every scene with a personPAI Ranking: the top 60% holds 95.7% of every scene with a personPAI Ranking: the top 70% holds 98.5% of every scene with a personPAI Ranking: the top 80% holds 99.7% of every scene with a personPAI Ranking: the top 90% holds 100.0% of every scene with a personPAI Ranking: the top 100% holds 100.0% of every scene with a person82%35%10%20%30%40%50%60%70%80%90%100%How far down the ranking you go
Both lines reach 100%, because the ranking reorders your archive rather than shrinking it.
The takeaway

Four bands in, you have reached most of the people in the archive. The rest of your footage is still there, in order, whenever you want it.

Table view
Depth of the rankingPAI RankingTemporal Order
Top 10%28.8%6.8%
Top 20%51.8%15.2%
Top 30%68.8%24.8%
Top 40%81.6%35.4%
Top 50%90.3%46.4%
Top 60%95.7%58.0%
Top 70%98.5%69.4%
Top 80%99.7%79.3%
Top 90%100.0%89.0%
Top 100%100.0%100.0%
4.92x

Moved to the Front of the Line

Say you have the time and the budget to review 5,000 scenes. That is one scene in every twenty.

Take the first 5,000 from our ranked queue and 1,525 of them have a person in them, which is 30 in every 100. Take the first 5,000 in the order the footage was recorded and 310 of them do, which is 6 in every 100. Across the whole archive the rate is 10 in every 100, so the recorded order actually starts below the archive's own average.

The same 5,000 scenes of work either way. Five times as many people in front of you at the end of it.

5x

Accelerated discovery, measured

From the same amount of review, the PAI Ranking reached 4.9 times as many scenes with people in them as Temporal Order.

PAI RankingTemporal Order, the order it came off the vehicle
0%10%20%30%40%PAI Ranking: 1,525 of the 5,000 scenes hold a person30.5%PAI Ranking1,525 of 5,000 scenesTemporal Order: 310 of the 5,000 scenes hold a person6.2%Temporal Order310 of 5,000 scenes
That is 1,525 scenes with people against 310, from the same 5,000 scenes of work.
The takeaway

WITNESS is not inference based, it’s simply computational, which means we can process your footage and make it more densely packed with meaningful scenes before your model or your human labeling process wastes time and money.

Table view
OrderScenes reviewedHolding a personShare
PAI Ranking5,0001,52530.5%
Temporal Order5,0003106.2%
Whole archive, for reference100,00010,36710.4%
4.10x

Our PAI Ranking System is more efficient.

Our ranking system makes your footage more densely packed with label-relevant scenes.

The chart before this one fixed the amount of work and counted what you got. This one turns the question around. It fixes what you want and counts the work.

Suppose your labeling pipeline needs 1,000 scenes with a person in them. Working down our queue you have them after opening 3,296 scenes. Working through the same footage in the order it was recorded you have them after opening 13,521. That is 10,225 fewer scenes opened for the same 1,000 found.

People-Dense Footage

The number down the left is how many scenes with a person in them you want. The bars are how many scenes you have to open to get that many. Shorter bars are better.

PAI RankingTemporal Order, the order it came off the vehicle
250PAI Ranking: 1,135 scenes opened to reach 2501,135Temporal Order: 3,810 scenes opened to reach 2503,810500PAI Ranking: 1,887 scenes opened to reach 5001,887Temporal Order: 7,477 scenes opened to reach 5007,4771,000PAI Ranking: 3,296 scenes opened to reach 1,0003,296Temporal Order: 13,521 scenes opened to reach 1,00013,5212,000PAI Ranking: 6,515 scenes opened to reach 2,0006,515Temporal Order: 24,279 scenes opened to reach 2,00024,2793,000PAI Ranking: 10,058 scenes opened to reach 3,00010,058Temporal Order: 33,991 scenes opened to reach 3,00033,9915,000PAI Ranking: 18,244 scenes opened to reach 5,00018,244Temporal Order: 51,639 scenes opened to reach 5,00051,639Scenes that must be opened to get there
To reach 1,000 scenes with people you open 3,296 instead of 13,521. Shorter bars are better.
The takeaway

Send us your entire archive or just enough to supply your labeling pipeline with the scenes it needs.

Table view
Scenes with people wantedPAI Ranking opensTemporal Order opensRatioLess footage opened
2501,1353,8103.36x70%
5001,8877,4773.96x75%
1,0003,29613,5214.10x76%
2,0006,51524,2793.73x73%
3,00010,05833,9913.38x70%
5,00018,24451,6392.83x65%
82% and 62%

People and objects, in the same ranking

Each band is 10,000 scenes, highest scored first. People are counted as scenes with a person in them. Objects are counted one by one.

Every chart so far has counted scenes with people. This one adds the other thing worth counting.

An object is one road user that somebody has to draw a box around: a car, a truck, a bus, a cyclist, a person. The archive holds 177,845 of them, spread across the 100,000 scenes.

The ranking is built to find people, so people is what it finds hardest. Objects follow along behind it, because a scene busy with people is usually busy with everything else as well. Both lines start far above an even spread and fall away together as you go down the list.

Scenes with peopleObjects to labelAn even spread, 10% in every band
0%8%16%24%32%People, Top tenth: 2,983 scenes, 28.8% of allObjects, Top tenth: 33,305 objects, 18.7% of allPeople, 2nd tenth: 2,391 scenes, 23.1% of allObjects, 2nd tenth: 29,546 objects, 16.6% of allPeople, 3rd tenth: 1,761 scenes, 17.0% of allObjects, 3rd tenth: 25,175 objects, 14.2% of allPeople, 4th tenth: 1,323 scenes, 12.8% of allObjects, 4th tenth: 21,484 objects, 12.1% of allPeople, 5th tenth: 905 scenes, 8.7% of allObjects, 5th tenth: 18,204 objects, 10.2% of allPeople, 6th tenth: 557 scenes, 5.4% of allObjects, 6th tenth: 15,386 objects, 8.7% of allPeople, 7th tenth: 291 scenes, 2.8% of allObjects, 7th tenth: 12,993 objects, 7.3% of allPeople, 8th tenth: 122 scenes, 1.2% of allObjects, 8th tenth: 11,222 objects, 6.3% of allPeople, 9th tenth: 30 scenes, 0.3% of allObjects, 9th tenth: 9,411 objects, 5.3% of allPeople, Last tenth: 4 scenes, 0.0% of allObjects, Last tenth: 1,119 objects, 0.6% of allAn even spread puts 10% in every tenthTop2nd3rd4th5th6th7th8th9thLastTenth of the ranking, best first
An unsorted archive would sit on the dashed line, with 10% of everything in every band.
The takeaway

One pass serves both. The ranking is built to find people, and the objects come with them.

Table view
Ranking band, 10% eachScenes with peopleShare of allAdmitted objectsShare of all
Top band2,98328.8%33,30518.7%
2nd band2,39123.1%29,54616.6%
3rd band1,76117.0%25,17514.2%
4th band1,32312.8%21,48412.1%
5th band9058.7%18,20410.2%
6th band5575.4%15,3868.7%
7th band2912.8%12,9937.3%
8th band1221.2%11,2226.3%
9th band300.3%9,4115.3%
Last band40.0%1,1190.6%

Measured on the Zenseact Open Dataset: 100,000 annotated scenes, published by Zenseact and free for anyone to download and re-run. The ranking reads lidar only. It never opens a label file and never runs a model, so you can run it on footage you haven't labeled yet.

The Archives

We tested WITNESS on three public driving archives.

We did not build our own test set. A test set we assemble ourselves proves nothing about the archive sitting on a customer's servers. All three archives below are published by other organizations, carry labels those organizations placed, and can be downloaded by anyone who wants to run the same measurement we ran.

NVIDIA PhysicalAI
Published by NVIDIA

An autonomous vehicle archive of twenty-second driving clips carrying radar, lidar and camera. Its labels are placed by machine rather than by people, so it shows how WITNESS behaves when the answer sheet is itself automated.

Driving clips1,350
Clip length20 seconds
Labels placed byMachine
nuScenes
Published by Motional

A city driving archive carrying radar, lidar and camera. Its labels are placed by human annotators, which makes it the strictest test we run: our order is graded against people rather than against another machine.

Annotated frames34,149
Distinct events639
Labels placed byPeople
Zenseact ZOD
Published by Zenseact

The largest archive we have run, carrying camera and lidar. It has no radar, which is why our first pass on it is a lidar pass. It is released under a licence that permits commercial use, so a customer can download it and repeat our measurement without asking anyone's permission.

Annotated frames100,000
LicenceCC BY-SA 4.0
Repeatable by youYes

The result: reviewing one scene in twenty, a team working in PAI Ranking order reaches 4.9 times as many scenes with people in them as a team reviewing the same driving footage in the order it was recorded. The figure is measured on the Zenseact Open Dataset.

NVIDIA and nuScenes ask that numbers measured on their datasets not be published. So the figure we publish is measured on ZOD, which is released under a licence that permits it.

Zenseact Open Dataset (ZOD), © 2022 Zenseact AB, licensed under CC BY-SA 4.0.

A vehicle on a highway at night beneath a star-filled sky and the Milky Way, its surroundings drawn as overlapping fields of measurement, with a city skyline ahead.
WITNESS measures the scene your vehicles already recorded
The process

Six Phases Create a Virtuous Cycle that Improves your Labeling Pipeline

The six phases of COSIMO WITNESS as a repeating cycle Survey, Find, Reduce, Trust, Validate, Prove, returning to Survey. REPEATS WITNESS Measured from the outside 01 SURVEY We measure where you stand before we touch anything. 02 FIND We analyze your entire archive, rank every scene, and move the scenes with the highest ranking to the start of your queue. 03 REDUCE We reach the same measured scene-ranking quality with far fewer labels. 04 TRUST We flag the automated, model-generated labels that are wrong. 05 VALIDATE We show you where your model is weakest, before anything is labeled. 06 PROVE We measure again and report what changed.
01SURVEY

We measure where you stand before we touch anything. What share of your existing labels are wrong, what the human work behind one accepted label costs you, and how much of the material worth labeling your current process finds.

02FIND

We analyze your entire archive, rank every scene, and move the scenes with the highest ranking to the start of your queue. No query to type, no starter examples to provide; the rare, edge-case scenes surface first.

03REDUCE

We reach the same scene-ranking quality with far fewer labels. WITNESS uses its PAI score, so ranking quality holds while the edge-case, label-relevant scenes are prioritized and presented for review ahead of the ordinary scenes.

04TRUST

We flag the automated, model-generated labels that are wrong. The industry assumes its labels are accurate, but published research shows otherwise. By flagging errors faster, WITNESS helps you troubleshoot your labeling pipeline.

05VALIDATE

A single mistake tells you little on its own. Read against your whole archive, that same mistake becomes one point in a larger pattern of where your model's understanding breaks down. WITNESS ranks that pattern and reveals the scenes your model handles worst, with no labels required.

06PROVE

We measure at the beginning and at the end of a single process. We report on what changed during our process. Both ends are measured the same way, so the improvement is a number your team can verify and validate.

Send us your archive

We rank it. You stop throwing away label-relevant scenes.

01
Send
Send us your entire multi-sensor driving footage archive, every sensor stream intact.
02
Rank
We process it and rank every scene by its label-relevance.
03
Return
You get the archive back, sequenced and organized in order of priority.

The result accelerates discovery by 2.5x to 2.7x on all label-relevant material, and by 4.2x to 4.9x on the scenes with people in them. Faster through the same archive, and no more setting aside scenes that were label-relevant all along.

Managed labeling

We also run the operation for you.

Send us the recordings with every sensor stream intact. We label them end to end, in house, and point the same instrument at our own work. You get the finished archive, both measurements, and the recipe to recompute them.

See how managed labeling works