Judgeval
The main entry point for interacting with the Judgment platform.
Judgeval connects to your Judgment project and gives you access to
SQL queries, evaluations, and datasets.
Credentials are resolved in order: explicit arguments first, then
environment variables JUDGMENT_API_KEY, JUDGMENT_ORG_ID, and
JUDGMENT_API_URL.
ValueError: If any required credential or project_name is missing.
Minimal setup (credentials from environment variables):
from judgeval import Judgeval
client = Judgeval(project_name="search-assistant")
Explicit credentials:
client = Judgeval(
project_name="search-assistant",
api_key="jdg_...",
organization_id="org_...",
)
Once initialized, use the evaluation and datasets
properties:
eval_runner = client.evaluation.create()
dataset = client.datasets.get(name="golden-set")
Attributes
evaluation?Any
Access evaluations for scoring examples with hosted or custom judges. Use .create() to get an Evaluation you can call .run() on. eval_runner = client.evaluation.create() results = eval_runner.run( examples=examples, scorers=["faithfulness", "answer_relevancy"], eval_run_name="nightly-eval", )
Anyoffline_tests?Any
Access offline tests: test configs and dataset-backed test runs. Use .create_config() to bind a dataset to judges, and .run() to execute a test run (optionally driving an agent entrypoint and asserting a pass condition). config = client.offline_tests.create_config( name="nightly-regression", dataset="golden-set", judges=["helpfulness"], ) result = client.offline_tests.run( test_config="nightly-regression", agent_function=my_agent, pass_condition_fn=lambda fields, scorers: all( s.error is None for s in scorers ), assert_test=True, )
Anydatasets?Any
Manage datasets of evaluation examples. Use .create(), .get(), or .list() to work with datasets. dataset = client.datasets.create( name="golden-set", schema={ "type": "object", "properties": { "input": {"type": "string"}, "expected_output": {"type": "string"}, }, }, examples=[ Example.create(input="What is 2+2?", expected_output="4"), ], )
Anyprompts?Any
Manage versioned prompt templates with tagging support. Use .create(), .get(), .tag(), or .list() to work with prompts. prompt = client.prompts.create( name="system-prompt", prompt="You are a helpful assistant for {{product}}.", tags=["v1"], ) compiled = prompt.compile(product="Acme Search")
Anyagent_judges?Any
Manage Agent Judges (prompt-based scorers) on the platform. Use .create() or .update() to create and update prompt-based Agent Judges. judge = client.agent_judges.create( name="helpfulness", prompt="Score the assistant's helpfulness from 0 to 1.", model="gpt-5.2", score_type="numeric", ) client.agent_judges.update( judge_id=judge.judge_id, prompt="Updated rubric prompt.", )
Anyquery()
Run a legacy JQL query, optionally narrowed by trace or session IDs.
Deprecated. Use sql() for new integrations, with SQL
predicates to narrow results. Existing JQL calls remain supported.
def query(query, *, limit=None, trace_ids=None, session_ids=None) -> 'JqlQueryResponse':
Parameters
query'QueryInput'
'QueryInput'limit?Optional[int]
Optional[int]Nonetrace_ids?Optional[Sequence[str]]
Optional[Sequence[str]]Nonesession_ids?Optional[Sequence[str]]
Optional[Sequence[str]]NoneReturns
'JqlQueryResponse'
discover_schema()
Return the SQL schema reference as Markdown.
Mirrors MCP discover_schema: published tables, column types and
descriptions, row semantics, examples, and query limits. Fetches the
server’s generated catalog using this client’s credentials. Contains
no project data and does not require a resolved project or public
query opt-in.
print(client.discover_schema())
def discover_schema() -> str:
Returns
str - The virtual schema reference as a Markdown string.
sql()
Run one read-only SQL SELECT for this organization and project.
Prefer this method for new read-only queries. The server derives tenant
scope from the client’s credentials and resolved project.
Call discover_schema() for supported tables and columns. Requires
viewer access and public SDK/API queries enabled for the organization.
Results are capped by the server at 1,000 rows and 5 MiB; exceeding either cap returns an error. Use SQL predicates and LIMIT to narrow results.
result = client.sql("SELECT count() AS run_count FROM telemetry.traces")
print(result["rows"])
def sql(sql_text) -> 'SqlResponse':
Parameters
sql_textstr
One SELECT against the virtual schema, at most 50,000 characters. Use SQL predicates to narrow the results.
strReturns
'SqlResponse' - A dictionary with catalog_version, columns (name, type, nullable),
rows (dictionaries keyed by column name), row_count, and
elapsed_ms. Integers outside JavaScript’s safe range arrive as
exact decimal strings.
present()
Run a legacy JQL chart or table query.
Deprecated. Use sql() for new queries and render its
rows as charts or tables in your application. SQL does not return a
JQL presentation frame. Existing presentation calls and their frame
responses remain supported.
def present(query, *, limit=None, trace_ids=None, session_ids=None) -> 'JqlPresentationResponse':
Parameters
query'QueryInput'
'QueryInput'limit?Optional[int]
Optional[int]Nonetrace_ids?Optional[Sequence[str]]
Optional[Sequence[str]]Nonesession_ids?Optional[Sequence[str]]
Optional[Sequence[str]]NoneReturns
'JqlPresentationResponse'
discover()
Discover project-scoped judges, fields, models, and related values.
Deprecated. Use discover_schema() to inspect
the SQL tables and columns, then sql() to query project values.
Schema discovery returns documentation, not project data. Existing
JQL discovery calls remain supported; SQL returns a different row schema.
def discover(kind, *, limit=None, trace_ids=None, session_ids=None, **options) -> 'JqlQueryResponse':
Parameters
kind'DiscoveryKind'
'DiscoveryKind'limit?Optional[int]
Optional[int]Nonetrace_ids?Optional[Sequence[str]]
Optional[Sequence[str]]Nonesession_ids?Optional[Sequence[str]]
Optional[Sequence[str]]Noneoptions?Any
Any{}Returns
'JqlQueryResponse'