Label a dataset
Review trace or session examples against a judge and create a separate dataset with human-expected labels.
Use Label Dataset to record the outputs a judge should return for examples in an existing dataset. Finishing creates a new dataset with the human labels; it leaves the source dataset unchanged.
Prerequisites
- A non-legacy dataset containing trace or session evidence
- A saved production judge with the matching evaluation scope: a trace judge for traces, or a session judge for sessions
- Permission to create a dataset
If you already reviewed judge results in Alignment, create a golden dataset from those reviews instead of labeling them again here.
Record the expected outputs
- Open Datasets, select the source dataset and version, and open Examples.
- Select Label Dataset.
- Choose a compatible judge, enter the New dataset name, and select Start labeling.
- Inspect each example’s trace or session. Record the expected output using the judge’s binary, categorical, or numeric controls. Use the evidence and rubric to decide the label, rather than copying an earlier judge result.
- Move through the examples and check the labeled count. When ready, select Create labeled dataset.
Only labeled examples are included in the new dataset. The footer shows how many unlabeled examples will be excluded, so finish reviewing every example first if you need the entire source version represented.
Verify the new dataset
Open the created dataset and inspect its example count, source fields, and expected-label column for the selected judge. Check a few labels against their trace or session evidence before using them as a test baseline.
The source dataset and selected source version remain available separately.
Troubleshooting
- If Label Dataset is absent, open Examples on a non-legacy dataset with a selected version and trace or session evidence.
- If the judge is missing from the picker, check its production version and whether its trace or session scope matches the dataset.
- If creation is disabled, record at least one label.
- If the new dataset has fewer examples, check whether some source examples were left unlabeled.
Next step
Run an offline test against the labeled dataset to compare judge revisions on the same evidence.