Skip to main content
Welcome to Labelbox Catalog, your command center for understanding, curating, and preparing unstructured data for your machine learning workflows. Before you can build a high-performing model, you need to deeply understand the data you’re working with. Catalog is designed to move you from raw data to curated, label-ready datasets with confidence and speed. Think of Catalog as an interactive, searchable index of all your training data. It provides the tools to not just see your data, but to interact with it, ask questions of it, and organize it in powerful ways.

What can you do with Catalog?

  • Visualize and explore your data to uncover insights: Instead of guessing what’s in your dataset, you can directly visualize it. Spot imbalances, identify outliers, find rare edge cases, and understand the distribution of your data before it ever touches a model. Use the gallery view for a visual survey, the list view for metadata analysis, and the analytics view to see statistical breakdowns.
  • Find specific data with powerful search and filtering: Move beyond simple filename searches. Catalog allows you to build complex queries to find the exact data you need. You can filter by a rich set of attributes including metadata, annotation-class, dataset, project, and even the content of the data itself using AI-powered search methods.
  • Curate and organize datasets for any workflow: Your raw data is just the beginning. Catalog helps you organize it for specific tasks. You can create static batches of data to send to a labeling project or define dynamic slices that automatically track specific subsets of your data over time, like “all images flagged for review.”
  • Take targeted action on your data: Finding data is only half the battle. Catalog is fully integrated with the rest of the Labelbox platform, allowing you to take immediate action on your findings. Select a group of data rows and, with a few clicks, you can add metadata, export them for analysis, or send them directly to a labeling project.
By providing a single, unified interface to explore and manage all your unstructured data, Catalog empowers you to make smarter, data-driven decisions throughout the entire model development lifecycle.

Key concepts

To effectively navigate and use Catalog, it’s important to understand its core components. These are the fundamental building blocks you’ll encounter as you explore and manage your data.