Scoring across interfaces
Distinguish dashboard evaluators from API, CLI, and MCP scoring modes.
The dashboard configures evaluator rules on a prompt. Management API, CLI, and MCP evaluation requests select a scoringMethod. These controls are different, so the same outputs can receive different results depending on the interface and scoring configuration.
Dashboard evaluators
Configure one or more evaluators and apply each to all or selected dataset cases. All applicable checks must pass for a case to pass. Initial dashboard setup requires an applicable evaluator for every case.
Rules include exact match, contains, regex, JSON schema, length, LLM score, and LLM checklist. Their options and passing rules are defined in the evaluator reference.
The first evaluation enables whitespace trimming and disables ignoring case. For reference billing, it accepts a trailing newline but rejects Billing.
Automation scoring modes
scoringMethod | Behavior |
|---|---|
expected | Compare the output with the case's reference; every selected case needs a nonblank reference |
llm_judge | Use the workspace judge model to assess the actual output against its reference; every selected case needs a nonblank reference |
human_review | Produce outputs for a person to review; results remain unscored until reviewed |
Initial baseline creation through automation requires expected scoring. The other automation modes apply to workflows such as later evaluations or human calibration.
For expected, when both values parse as JSON, comparison ignores object key order and preserves array order and value types. Otherwise, text comparison trims leading and trailing whitespace and ignores case. For reference billing, Billing therefore passes; billing. and Team: billing fail.
The llm_judge mode uses a built-in rubric against the reference output. It is separate from the dashboard's configurable LLM score and checklist evaluators. Configure the workspace judge model before using it.
Management requests do not accept the dashboard's evaluator definitions as scoringMethod. Use the API operation schemas for accepted fields.
Moving between interfaces
Check the scoring mode, reference requirements, case selection, and matching options before comparing results from different interfaces. The ticket-routing examples share a dataset and prompt task, but their dashboard and automation scoring rules are deliberately documented separately.
For the complete automation sequence, follow the CLI walkthrough or management API walkthrough.