Author: )
Hey there! , and I recently gave an in-person workshop on How to Unlock More in Self-Driving Datasets. I’m super excited to dive into this with you today and bring it to everyone! I highly recommend checking out my video where I dive into detail about the workshop, and to check out the
The self-driving revolution is in full swing, and the question remains: Who will break the threshold to release self-driving cars into the world safely? A combination of hardware, software, and strategy is shaping this landscape.
💪 The Power of GPUs
was a nightmare. Today, it takes hours, not days, thanks to optimization and community support. These advancements have streamlined workflows and allowed for faster experimentation.
🌐 Data at Scale
, Voxel 51’s open-source dataset management tool, come in.
🧠 Top-Tier Talent Behind the Scenes
The self-driving field attracts the best engineers and researchers. Companies like Waymo, Wayve, and Tesla are leading the charge, but much of their work remains behind closed doors for competitive reasons. Recently, however, we’ve seen a shift, with more research being published. This openness is giving us a peek into what makes their labs so innovative.
🎲 The Big Gamble
The strategies for self-driving success are as varied as the companies pursuing them:
- Wayve: Aims for an end-to-end solution, teaching cars to drive anywhere by building world models.
- Waymo: Focuses on mastering individual cities with detailed maps, making their vehicles highly efficient in those areas.
- Tesla: Stands apart by relying solely on image-based systems, forgoing LiDAR entirely—a bold but controversial approach.
Who will win? It’s anyone’s guess. The competition is a massive gamble, and it’s thrilling to watch.
🛠️ Beginner Techniques: Curation, Digitization, and Dataset Management
Let’s get practical. Whether you’re a beginner or an expert, organizing self-driving data is the foundation. Two major challenges you’ll face:
Unstructured Data: Multi-camera systems, different sensors, varying frame rates—it’s all a big jumble.
Scale: Even hobbyists deal with massive datasets, so efficient organization is key.
This is where FiftyOne shines. With FiftyOne, you can:
- Load diverse data (images, videos, radar, LiDAR) seamlessly.
- Visualize, clean, and curate datasets.
- Debug datasets, find gaps, and evaluate model performance.
🔧 Building a Grouped Dataset
. The result? A well-organized dataset ready for training.
📂 FiftyOne: Your Dataset Debugger
is a great tool here! SAM2 and more are included in the
While pretrained models help with recognizing objects, embeddings help us understand how similar or dissimilar various samples in our dataset are to one another. Using both 2D and 3D embedding models, we can create a map of our dataset, highlighting clusters of similar samples and pinpointing where there may be gaps. This is easy with the
Once we’ve generated our embeddings, we can head to the embeddings panel in FiftyOne to view the results. It’s as simple as hitting the plus button, selecting the "brain key," and using the embeddings visualization. You can also color-code the points based on different metadata stored in the dataset, like scene tokens.
As we dive into the embeddings grid, we get to see the relationships between different samples. You might notice, for example, that one cluster is all related to nighttime scenes, while another is more typical of daytime driving scenarios. These groupings help us understand what’s going on in the data and spot potential outliers that might be worth investigating.
🕵️ The Power of Embeddings for Dataset Curation
So far, we’ve covered some of the most advanced techniques available to enhance your self-driving datasets. But the experts in the field are pushing things even further.
One of the biggest hurdles in self-driving technology is time. Training models takes time, and testing those models in real-world scenarios requires a lot of trial and error. What if you could eliminate this cycle and simulate scenarios in a controlled environment? That’s where simulation comes in.
With tools like for code snippets and examples to help you build your first grouped dataset.
SOCIAL SHARE CARD GENERATOR