Skip to content

Results & Scoring

When an evaluation job completes, EvalHub populates a results section on the job resource. This page explains exactly how the test values are calculated — both the per-benchmark pass/fail (results.benchmarks[].test) and the job-level aggregate (results.test) — and which configuration and request overrides change the outcome.

The results section carries two distinct test concepts:

FieldTypeMeaning
results.benchmarks[].testper-benchmarkPass/fail for a single benchmark, based on its chosen primary metric vs. its threshold.
results.testjob-levelOverall pass/fail for the whole job, a weighted average of every benchmark’s primary score vs. the job threshold.
{
"results": {
"test": {
"score": 0.72,
"threshold": 0.45,
"pass": true
},
"benchmarks": [
{
"id": "arc_easy",
"provider_id": "lm_evaluation_harness",
"benchmark_index": 0,
"metrics": { "acc": 0.81, "acc_norm": 0.78 },
"test": {
"primary_score": 0.78,
"primary_score_metric": "acc_norm",
"threshold": 0.25,
"pass": true
}
},
{
"id": "toxigen",
"provider_id": "lm_evaluation_harness",
"benchmark_index": 1,
"metrics": { "acc": 0.60 },
"test": {
"primary_score": 0.60,
"primary_score_metric": "acc",
"threshold": 0.70,
"pass": false
}
}
]
}
}
  • metrics — the full map of metric values the adapter reported for that benchmark (passthrough).
  • benchmarks[].test — the derived pass/fail for that benchmark (see below).
  • test — the derived job-level result (see below).

results.benchmarks[].test is recomputed on every benchmark status event and stored once the benchmark reaches a terminal state. The steps are:

  1. Resolve the effective benchmark config. If the job references a collection, the benchmark configuration is taken from the collection (with any per-benchmark job overrides merged in); otherwise it comes from the job’s inline benchmarks.

  2. Choose the primary metric. Use primary_score.metric from the job/collection benchmark config. If that is unset or empty, fall back to the provider’s default primary_score.metric for that benchmark.

  3. Look up the value. Read metrics[primary_metric] from the values the adapter reported. The value must be numeric and is cast to a float (see Non-numeric metrics below).

    • If the primary metric is not present in the reported metrics, no test block is produced for that benchmark (the field is omitted).
    • If the value is present but not numeric, the same thing happens — the test block is omitted.
  4. Choose the threshold. Use pass_criteria.threshold from the job/collection benchmark config. If unset, fall back to the provider’s default pass_criteria.threshold for that benchmark.

    • If neither defines a threshold, no test block is produced.
  5. Decide pass/fail.

    pass = primary_score >= threshold # default
    pass = primary_score <= threshold (if lower_is_better)

The resulting object is:

{
"primary_score": 0.78, // the reported value of the chosen metric
"primary_score_metric": "acc_norm",
"threshold": 0.25,
"pass": true
}

Scoring is purely arithmetic: the primary score must be a number, and the job-level aggregate is a weighted mean of numbers. EvalHub has no notion of combining categorical, string, or boolean outcomes into a test result.

When EvalHub reads the primary metric’s value, it casts it to a float and accepts only these types:

AcceptedRejected
float64, float32strings (including numeric-looking strings like "0.78")
int, int32, int64booleans (true / false)
arrays and objects
null / absent

JSON numbers arrive as float64, so any adapter that reports a plain numeric metric works without special handling.

What happens when the primary metric is non-numeric

Section titled “What happens when the primary metric is non-numeric”

The effect cascades from the single benchmark up to the whole job:

  1. Benchmark level. If the chosen primary_score.metric value is non-numeric, the benchmark’s test computation fails the cast and produces no test block. The failure is logged server-side, but the API response simply omits benchmarks[].test for that benchmark. The raw value is still preserved under benchmarks[].metrics.
  2. Job level. Any benchmark with no test block is skipped in the weighted average — it contributes to neither the weighted sum nor the total weight, so it does not drag the score up or down; it is simply absent from the calculation.
  3. Whole job. If no benchmark yields a numeric primary score (so no weights accumulate), the job score cannot be computed and results.test is omitted entirely.
  • Report numeric metrics for scoring. If an adapter naturally emits a categorical or textual result, convert it to a numeric metric (for example a 0/1 pass indicator, a rate, or a normalized score) and point primary_score.metric at that numeric field.
  • Keep the rich output as passthrough. Non-numeric values (labels, category breakdowns, free-text notes) can still be reported and are retained under benchmarks[].metrics — they are just ignored by scoring, not discarded.

results.test is computed once, when the overall job reaches the completed state. It is a weighted arithmetic mean of every benchmark’s primary score:

Σ ( weightᵢ × scoreᵢ )
job.score = ─────────────────────────────
Σ weightᵢ
pass = job.score >= job.threshold

Where, for each benchmark i:

  • scoreᵢ is that benchmark’s test.primary_score.
    • If the benchmark’s metric is lower_is_better, the contribution is inverted to 1 - primary_score so that a higher aggregate always means “better”. (Because of this inversion, the job-level comparison is always >=.)
  • weightᵢ is the benchmark’s configured weight. A missing or 0 weight is treated as 1.

Benchmarks that have no test block (for example, a missing primary metric or threshold) are skipped — they contribute to neither the numerator nor the denominator. If no weights accumulate at all (Σ weightᵢ == 0), the job score is not computed and results.test is omitted.

Two benchmarks, one higher-is-better and one lower-is-better:

Benchmarkprimary_scorelower_is_betterweightcontribution
arc_easy0.78no22 × 0.78 = 1.56
toxicity0.10yes11 × (1 − 0.10) = 0.90
job.score = (1.56 + 0.90) / (2 + 1) = 2.46 / 3 = 0.82

With a job threshold of 0.45, 0.82 >= 0.45 → pass: true.

User overrides that affect the test results

Section titled “User overrides that affect the test results”

The scoring inputs — primary metric, threshold, weight, and lower_is_better — can each be set in several places. When more than one is present, the most specific wins.

Selects which metric becomes primary_score for a benchmark.

PrioritySource
1 (highest)primary_score.metric on the job/collection benchmark config
2 (fallback)The provider’s default primary_score.metric for that benchmark

If neither defines a metric, the benchmark gets no test block.

Determines whether a single benchmark passes.

PrioritySource
1 (highest)pass_criteria.threshold on the job/collection benchmark config
2 (fallback)The provider’s default pass_criteria.threshold for that benchmark

If neither defines a threshold, the benchmark gets no test block (and is therefore excluded from the job-level average).

Controls each benchmark’s influence on the job-level average.

PrioritySource
1 (highest)weight on the job/collection benchmark config
2 (fallback)Default 1 (a 0 weight is also treated as 1)

Determines whether the whole job passes.

PrioritySourceExample use case
1 (highest)pass_criteria.threshold on the job request”For just this run, use a stricter bar of 0.9”
2pass_criteria.threshold on the collection definitionThe collection’s default bar
3 (fallback)Hard-coded default 0.5Neither job nor collection defines one

Set on a benchmark’s primary_score. It changes two things:

  • Per benchmark: the pass comparison flips to primary_score <= threshold.
  • Job level: the benchmark’s contribution is inverted to 1 - primary_score before it is weighted and averaged.

Use it for metrics where a smaller number is better (for example attack_success_rate or a toxicity rate).

Example: overriding scoring on a job request

Section titled “Example: overriding scoring on a job request”
POST /api/v1/evaluations/jobs
{
"name": "stricter-safety-run",
"model": { "...": "..." },
"pass_criteria": { "threshold": 0.9 },
"collection": {
"id": "safety-and-fairness-v1",
"benchmarks": [
{
"id": "toxigen",
"provider_id": "lm_evaluation_harness",
"weight": 3,
"primary_score": { "metric": "acc", "lower_is_better": false },
"pass_criteria": { "threshold": 0.7 }
}
]
}
}

Here the job raises the overall bar to 0.9, and re-weights/re-thresholds the toxigen benchmark just for this run, without changing the shared collection.

The test block is deliberately omitted (rather than showing a misleading 0) when:

  • the chosen primary metric is not present in the adapter’s reported metrics, or
  • the primary metric value is non-numeric (see Non-numeric metrics), or
  • no threshold can be resolved (neither the benchmark config nor the provider defines one), or
  • (job level) no benchmarks contributed a valid weight/score.
  • Collections — where weights, primary scores, and thresholds are configured, and the list of built-in collections and their thresholds.
  • Evaluation Event Violations — how a failing benchmark test triggers a threshold-violation notification.
  • Server API — the full evaluation job and results schema.