---
title: "Label a dataset"
description: "Review trace or session examples against a judge and create a separate dataset with human-expected labels."
sidebar:
  label: "Label a dataset"
seo:
  title: "Label a Dataset | Judgment How-to"
  description: "Choose a dataset version and compatible judge, record human labels, and create a labeled dataset for offline evaluation."
---

Use **Label Dataset** to record the outputs a judge should return for examples
in an existing dataset. Finishing creates a new dataset with the human labels;
it leaves the source dataset unchanged.

## Prerequisites

- A non-legacy dataset containing trace or session evidence
- A saved production judge with the matching evaluation scope: a trace judge
  for traces, or a session judge for sessions
- Permission to create a dataset

If you already reviewed judge results in Alignment, [create a golden
dataset](/documentation/tests/judge-calibration#create-a-labeled-dataset) from
those reviews instead of labeling them again here.

## Record the expected outputs

1. Open **Datasets**, select the source dataset and version, and open **Examples**.
2. Select **Label Dataset**.
3. Choose a compatible judge, enter the **New dataset name**, and select
   **Start labeling**.
4. Inspect each example's trace or session. Record the expected output using
   the judge's binary, categorical, or numeric controls. Use the evidence and
   rubric to decide the label, rather than copying an earlier judge result.
5. Move through the examples and check the labeled count. When ready, select
   **Create labeled dataset**.

Only labeled examples are included in the new dataset. The footer shows how
many unlabeled examples will be excluded, so finish reviewing every example
first if you need the entire source version represented.

## Verify the new dataset

Open the created dataset and inspect its example count, source fields, and
expected-label column for the selected judge. Check a few labels against their
trace or session evidence before using them as a test baseline.

The source dataset and selected source version remain available separately.

## Troubleshooting

- If **Label Dataset** is absent, open **Examples** on a non-legacy dataset
  with a selected version and trace or session evidence.
- If the judge is missing from the picker, check its production version and
  whether its trace or session scope matches the dataset.
- If creation is disabled, record at least one label.
- If the new dataset has fewer examples, check whether some source examples
  were left unlabeled.

## Next step

[Run an offline test](/documentation/tests/offline-tests) against the labeled
dataset to compare judge revisions on the same evidence.
