Skip to content
Judgment Labs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Judgeval

The main entry point for interacting with the Judgment platform.

Judgeval connects to your Judgment project and gives you access to SQL queries, evaluations, and datasets.

Credentials are resolved in order: explicit arguments first, then environment variables JUDGMENT_API_KEY, JUDGMENT_ORG_ID, and JUDGMENT_API_URL.

ValueError: If any required credential or project_name is missing.

Minimal setup (credentials from environment variables):

from judgeval import Judgeval

client = Judgeval(project_name="search-assistant")

Explicit credentials:

client = Judgeval(
    project_name="search-assistant",
    api_key="jdg_...",
    organization_id="org_...",
)

Once initialized, use the evaluation and datasets properties:

eval_runner = client.evaluation.create()
dataset = client.datasets.get(name="golden-set")

Attributes

PropType
evaluation?Any

Access evaluations for scoring examples with hosted or custom judges. Use .create() to get an Evaluation you can call .run() on. eval_runner = client.evaluation.create() results = eval_runner.run( examples=examples, scorers=["faithfulness", "answer_relevancy"], eval_run_name="nightly-eval", )

TypeAny
offline_tests?Any

Access offline tests: test configs and dataset-backed test runs. Use .create_config() to bind a dataset to judges, and .run() to execute a test run (optionally driving an agent entrypoint and asserting a pass condition). config = client.offline_tests.create_config( name="nightly-regression", dataset="golden-set", judges=["helpfulness"], ) result = client.offline_tests.run( test_config="nightly-regression", agent_function=my_agent, pass_condition_fn=lambda fields, scorers: all( s.error is None for s in scorers ), assert_test=True, )

TypeAny
datasets?Any

Manage datasets of evaluation examples. Use .create(), .get(), or .list() to work with datasets. dataset = client.datasets.create( name="golden-set", schema={ "type": "object", "properties": { "input": {"type": "string"}, "expected_output": {"type": "string"}, }, }, examples=[ Example.create(input="What is 2+2?", expected_output="4"), ], )

TypeAny
prompts?Any

Manage versioned prompt templates with tagging support. Use .create(), .get(), .tag(), or .list() to work with prompts. prompt = client.prompts.create( name="system-prompt", prompt="You are a helpful assistant for {{product}}.", tags=["v1"], ) compiled = prompt.compile(product="Acme Search")

TypeAny
agent_judges?Any

Manage Agent Judges (prompt-based scorers) on the platform. Use .create() or .update() to create and update prompt-based Agent Judges. judge = client.agent_judges.create( name="helpfulness", prompt="Score the assistant's helpfulness from 0 to 1.", model="gpt-5.2", score_type="numeric", ) client.agent_judges.update( judge_id=judge.judge_id, prompt="Updated rubric prompt.", )

TypeAny

query()

Run a legacy JQL query, optionally narrowed by trace or session IDs.

Deprecated. Use sql() for new integrations, with SQL predicates to narrow results. Existing JQL calls remain supported.

def query(query, *, limit=None, trace_ids=None, session_ids=None) -> 'JqlQueryResponse':

Parameters

PropType
query'QueryInput'
Type'QueryInput'
limit?Optional[int]
TypeOptional[int]
DefaultNone
trace_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone
session_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone

Returns

'JqlQueryResponse'


discover_schema()

Return the SQL schema reference as Markdown.

Mirrors MCP discover_schema: published tables, column types and descriptions, row semantics, examples, and query limits. Fetches the server’s generated catalog using this client’s credentials. Contains no project data and does not require a resolved project or public query opt-in.

print(client.discover_schema())
def discover_schema() -> str:

Returns

str - The virtual schema reference as a Markdown string.


sql()

Run one read-only SQL SELECT for this organization and project.

Prefer this method for new read-only queries. The server derives tenant scope from the client’s credentials and resolved project. Call discover_schema() for supported tables and columns. Requires viewer access and public SDK/API queries enabled for the organization.

Results are capped by the server at 1,000 rows and 5 MiB; exceeding either cap returns an error. Use SQL predicates and LIMIT to narrow results.

result = client.sql("SELECT count() AS run_count FROM telemetry.traces")
print(result["rows"])
def sql(sql_text) -> 'SqlResponse':

Parameters

PropType
sql_textstr

One SELECT against the virtual schema, at most 50,000 characters. Use SQL predicates to narrow the results.

Typestr

Returns

'SqlResponse' - A dictionary with catalog_version, columns (name, type, nullable), rows (dictionaries keyed by column name), row_count, and elapsed_ms. Integers outside JavaScript’s safe range arrive as exact decimal strings.


present()

Run a legacy JQL chart or table query.

Deprecated. Use sql() for new queries and render its rows as charts or tables in your application. SQL does not return a JQL presentation frame. Existing presentation calls and their frame responses remain supported.

def present(query, *, limit=None, trace_ids=None, session_ids=None) -> 'JqlPresentationResponse':

Parameters

PropType
query'QueryInput'
Type'QueryInput'
limit?Optional[int]
TypeOptional[int]
DefaultNone
trace_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone
session_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone

Returns

'JqlPresentationResponse'


discover()

Discover project-scoped judges, fields, models, and related values.

Deprecated. Use discover_schema() to inspect the SQL tables and columns, then sql() to query project values. Schema discovery returns documentation, not project data. Existing JQL discovery calls remain supported; SQL returns a different row schema.

def discover(kind, *, limit=None, trace_ids=None, session_ids=None, **options) -> 'JqlQueryResponse':

Parameters

PropType
kind'DiscoveryKind'
Type'DiscoveryKind'
limit?Optional[int]
TypeOptional[int]
DefaultNone
trace_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone
session_ids?Optional[Sequence[str]]
TypeOptional[Sequence[str]]
DefaultNone
options?Any
TypeAny
Default{}

Returns

'JqlQueryResponse'

Was this page helpful?