Skip to main content
A Studio dataset is a different object from a curated dataset. A curated dataset is a frozen union of batch runs, downloaded as one JSON. A Studio dataset is what sits behind Dataset Studio in the app: it carries its own taxonomy schemas, every event gets classified against them, and the Explore rail counts how many events fall under each value.

Two calls

A batch that ran with an output_schema has already classified its events, so creating a Studio from it adopts that schema as a taxonomy — the dimensions are there immediately, with no second analysis. Calling add_taxonomy with the same schema afterwards is rejected as a duplicate, which is correct: the work is already done. create_from_batch takes a batch id, an analyze() result, or a list of either — the batches become the dataset’s buckets directly, with no folder in between. You can also build one from folders with client.studio.create(title=..., folder_ids=[...]). add_taxonomy is for classifying a dataset against additional dimensions — fields the source batch never produced, or a dataset built from folders rather than from a schema-carrying batch:
It takes the same schema forms analyze(output_schema=...) does: a Pydantic model, a list of field dicts, or a schema dict. It runs as a real structured-ask batch over the dataset’s videos, so it is asynchronous — results appear as it completes.
Put every field of one taxonomy in a single add_taxonomy call. Each call lands its results in its own bucket of events, and the rail reconstructs cross-taxonomy combinations by time overlap. That is fine for independent dimensions and wrong for facets of one observation — a clip with a yellow line on the left and a white one on the right would produce both cross-products. One call keeps a combination something observed rather than inferred.

Coverage

The Explore rail counts one value at a time, conditioned on what is already selected. coverage() returns the whole joint distribution in one pass — the “which combinations do we actually have examples of” table:
Group on a subset with coverage(dataset_id, ["location", "lane_line_type"]). A combination with no examples is absent from the table rather than a zero row — usually the answer you were looking for. A repeated taxonomy field is expanded, so an event counts under each of its values — with two repeated fields that is a cross-product, which is right for “which combinations have examples” and wrong for “how many events”, so such a table will not sum to the event count. Counts are of what each event owns. A value inherited from a time-overlapping event of another taxonomy can satisfy a filter in the rail, but counting it here would let two sibling events report the same combination twice and overstate coverage. Pass include_inherited=True for the rail’s broader matching. For the raw table, client.studio.facet_rows(dataset_id).to_dataframe() gives one row per event with a column per taxonomy field. Values inherited from a time-overlapping event of another taxonomy are included by default; pass include_inherited=False to keep only what each event owns. coverage() and to_dataframe() need pandas.

From the app

The batch viewer’s actions menu has Create Studio, which does the same thing as create_from_batch for the batch you are looking at and opens the new dataset.