A Studio dataset is a different object from a curated dataset.
A curated dataset is a frozen union of batch runs, downloaded as one JSON. A
Studio dataset is what sits behind Dataset Studio in the app: it carries its own
taxonomy schemas, every event gets classified against them, and the Explore rail
counts how many events fall under each value.
Two calls
A batch that ran with an output_schema has already classified its events, so
creating a Studio from it adopts that schema as a taxonomy — the dimensions
are there immediately, with no second analysis. Calling add_taxonomy with the
same schema afterwards is rejected as a duplicate, which is correct: the work
is already done.
create_from_batch takes a batch id, an analyze() result, or a list of
either — the batches become the dataset’s buckets directly, with no folder in
between. You can also build one from folders with
client.studio.create(title=..., folder_ids=[...]).
Multi-query batches
A multi-query run (a structured-ask batch with several queries) is one parent
batch with one child batch per query. Pass the parent’s id and the Studio
gets one bucket per query, labelled the way the batch viewer’s query rail
labels them:
The parent itself becomes the dataset’s parent, so nothing is re-derived or
duplicated. A query that failed is left out (it has no events to bucket); a
query that is still running is an error, since the Studio would miss its events.
Passing a single child id instead builds a Studio of that one query only. Each
query that ran with its own output_schema seeds its own taxonomy, pointed at
its own bucket.
add_taxonomy is for classifying a dataset against additional dimensions —
fields the source batch never produced, or a dataset built from folders rather
than from a schema-carrying batch:
It takes the same schema forms analyze(output_schema=...) does: a Pydantic
model, a list of field dicts, or a schema dict. It runs as a real
structured-ask batch over the dataset’s videos, so it is asynchronous — results
appear as it completes.
Put every field of one taxonomy in a single add_taxonomy call. Each call
lands its results in its own bucket of events, and the rail reconstructs
cross-taxonomy combinations by time overlap. That is fine for independent
dimensions and wrong for facets of one observation — a clip with a yellow line
on the left and a white one on the right would produce both cross-products.
One call keeps a combination something observed rather than inferred.
Coverage
The Explore rail counts one value at a time, conditioned on what is already
selected. coverage() returns the whole joint distribution in one pass — the
“which combinations do we actually have examples of” table:
Group on a subset with coverage(dataset_id, ["location", "lane_line_type"]).
A combination with no examples is absent from the table rather than a zero
row — usually the answer you were looking for.
A repeated taxonomy field is expanded, so an event counts under each of its
values — with two repeated fields that is a cross-product, which is right for
“which combinations have examples” and wrong for “how many events”, so such a
table will not sum to the event count.
Counts are of what each event owns. A value inherited from a
time-overlapping event of another taxonomy can satisfy a filter in the rail,
but counting it here would let two sibling events report the same combination
twice and overstate coverage. Pass include_inherited=True for the rail’s
broader matching.
For the raw table, client.studio.facet_rows(dataset_id).to_dataframe() gives
one row per event with a column per taxonomy field. Values inherited from a
time-overlapping event of another taxonomy are included by default; pass
include_inherited=False to keep only what each event owns.
coverage() and to_dataframe() need pandas.
From the app
The batch viewer’s actions menu has Create Studio, which does the same
thing as create_from_batch for the batch you are looking at and opens the new
dataset.