Skip to content
Judgment Labs
Esc
↑↓navigate↵open⌘Jpreview
On this page

EvaluatorRunner

Abstract base for evaluation runners.

PropType
S?Any
TypeAny
DefaultTypeVar('S', str, Judge)
ExperimentRunItem?Any
TypeAny
DefaultDict[str, Any]

Abstract base for evaluation runners.

Concrete implementations handle either hosted (server-side) or local (in-process) scorer execution. The generic parameter S is str for hosted scorers or Judge for local scorers.

__init__()

def __init__(client, project_id, project_name):

Parameters

PropType
clientJudgmentSyncClient
TypeJudgmentSyncClient
project_idOptional[str]
TypeOptional[str]
project_namestr
Typestr

run()

Execute an evaluation run and return results.

def run(examples, scorers, eval_run_name, assert_test=False, timeout_seconds=300) -> typing.List:

Parameters

PropType
examplesList[Example]

Examples to evaluate.

TypeList[Example]
scorersList[S]

Scorers to run (strings or Judge instances).

TypeList[S]
eval_run_namestr

Name for this evaluation run.

Typestr
assert_test?bool

Deprecated and ignored by the current evaluation result payload.

Typebool
DefaultFalse
timeout_seconds?int

Maximum time to wait for results.

Typeint
Default300

Returns

typing.List - A list of ScoringResult objects, one per example.


_binary_label()

def _binary_label(value) -> str:

Parameters

PropType
valuebool
Typebool

Returns

str


_scorer_value()

def _scorer_value(scorer_dict) -> str | float | None:

Parameters

PropType
scorer_dictMapping[str, Any]
TypeMapping[str, Any]

Returns

str | float | None

Was this page helpful?