EvaluatorRunner
Abstract base for evaluation runners.
PropType
S?Any
Type
AnyDefault
TypeVar('S', str, Judge)ExperimentRunItem?Any
Type
AnyDefault
Dict[str, Any]Abstract base for evaluation runners.
Concrete implementations handle either hosted (server-side) or local
(in-process) scorer execution. The generic parameter S is str
for hosted scorers or Judge for local scorers.
__init__()
def __init__(client, project_id, project_name):
Parameters
PropType
clientJudgmentSyncClient
Type
JudgmentSyncClientproject_idOptional[str]
Type
Optional[str]project_namestr
Type
strrun()
Execute an evaluation run and return results.
def run(examples, scorers, eval_run_name, assert_test=False, timeout_seconds=300) -> typing.List:
Parameters
PropType
examplesList[Example]
Examples to evaluate.
Type
List[Example]scorersList[S]
Scorers to run (strings or Judge instances).
Type
List[S]eval_run_namestr
Name for this evaluation run.
Type
strassert_test?bool
Deprecated and ignored by the current evaluation result payload.
Type
boolDefault
Falsetimeout_seconds?int
Maximum time to wait for results.
Type
intDefault
300Returns
typing.List - A list of ScoringResult objects, one per example.
_binary_label()
def _binary_label(value) -> str:
Parameters
PropType
valuebool
Type
boolReturns
str
_scorer_value()
def _scorer_value(scorer_dict) -> str | float | None:
Parameters
PropType
scorer_dictMapping[str, Any]
Type
Mapping[str, Any]Returns
str | float | None