Eval Hub (1.0.3)

Download OpenAPI specification:

License: Apache 2.0

API REST server for evaluation backend orchestration

Evaluations

Evaluation job management endpoints

Create Evaluation

Create and execute evaluation request using the simplified benchmark schema.

Request Body schema: application/json
required
One of
name
required
string

The evaluation job name.

description
string

The evaluation job description.

tags
Array of strings

The evaluation job tags.

required
object (ModelRef)

The model to evaluate, or the model that was used to generate the pre-recorded data.

required
Array of objects (EvaluationBenchmarkConfig)

The evaluation benchmarks to run.

object (PassCriteriaWithDefault)

The overall pass criteria for the evaluation job.

object (ExperimentConfig)

The MLFlow experiment configuration. When provided, the evaluation job will be tracked in MLFlow.

object (EvaluationExports)

Optional exports configuration for the evaluation job. When provided, the evaluation job results will be exported to the specified location.

object (BenchmarkHardwareConfig)

Optional evaluation-level hardware override for Kubernetes-backed jobs. Applied as a fallback for every benchmark that does not specify its own hardware_config. Prefer per-benchmark hardware_config when individual benchmarks need different hardware. When neither is set, the deprecated top-level queue field may still supply a scheduling queue.

object (QueueConfig)
Deprecated

Deprecated. Prefer benchmark.hardware_config or evaluation-level hardware_config instead. This field will be discontinued in a future release. Retained for backward compatibility: when neither benchmark.hardware_config nor evaluation hardware_config is set, this queue is applied to scheduled jobs. When a benchmark references a queue-backed HardwareProfile (or sets hardware_config.queue), that configuration takes precedence over this field.

object

Custom request data. This can be used for user specific job data.

Responses

Request samples

Content type
application/json
Example
{
  • "name": "granite-3.1-8b-safety-eval",
  • "description": "Safety and reasoning evaluation for Granite 3.1 8B Instruct",
  • "tags": [
    ],
  • "model": {},
  • "benchmarks": [
    ],
  • "pass_criteria": {
    }
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "status": {
    },
  • "name": "granite-3.1-8b-safety-eval",
  • "description": "Safety and reasoning evaluation for Granite 3.1 8B Instruct",
  • "tags": [
    ],
  • "model": {},
  • "benchmarks": [
    ],
  • "pass_criteria": {
    }
}

List Evaluations

List all evaluation requests.

query Parameters
limit
integer (Limit) [ 1 .. 100 ]
Default: 50

Maximum number of evaluations to return

offset
integer (Offset) >= 0
Default: 0

Offset for pagination

status
string (Status Filter)

Filter by status

name
string (Name)

Name to search for

tags
string (Tags)

Tags to search for

experiment_id
string (Experiment Id)

Filter by MLflow experiment ID

collection_id
string (Collection Id)

Filter by the ID of the collection the job was created from. Maps to the collection.id field on the job record. Returns all jobs for the collection.

Responses

Response samples

Content type
application/json
{
  • "first": {
    },
  • "next": {
    },
  • "limit": 50,
  • "total_count": 73,
  • "items": [
    ]
}

Get Evaluation

Returns the evaluation job resource with the current status and results.

path Parameters
id
required
string (Id)

Responses

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "status": {
    },
  • "results": {
    },
  • "name": "granite-3.1-8b-safety-eval",
  • "description": "Safety and reasoning evaluation for Granite 3.1 8B Instruct",
  • "tags": [
    ],
  • "model": {},
  • "benchmarks": [
    ],
  • "pass_criteria": {
    }
}

Cancel Evaluation

Cancel a running evaluation.

path Parameters
id
required
string (Id)
query Parameters
hard_delete
boolean (Hard Delete)
Default: false

If true, delete the evaluation job permanently so that GET /api/v1/evaluations/jobs/{id} will return a 404.

Responses

Response samples

Content type
application/json
{
  • "message": "The field 'state' is not valid.",
  • "message_code": "invalid_value",
  • "trace": "b12692e1-8582-4628-88ca-7a13fefb73e2"
}

Get Evaluation Job Logs

Returns plain-text workload logs for all benchmarks in an evaluation job.

Kubernetes runtime: adapter container stdout/stderr via the Kubernetes API. Local runtime: contents of each benchmark's jobrun.log file under /tmp/evalhub-jobs/{job_id}/{benchmark_index}/{provider_id}/{benchmark_id}/. Logs are fetched on demand from the active runtime. Distinct from logs_path on benchmark results, which refers to adapter-written artifact files.

path Parameters
id
required
string (Id)
query Parameters
-1 (integer) or integer
Default: 1000

Maximum number of log lines to return per benchmark. The response concatenates one section per benchmark; each section is capped independently, not as a total across the full response. Use -1 to request all available log lines (no tail limit). Valid values: 1–10000, or -1 for all lines.

timestamps
boolean
Default: false

Include Kubernetes log timestamps

since_seconds
integer >= 1

Only return logs newer than this many seconds

Responses

Response samples

Content type
text/plain
=== pod=a1b2c3d4-405ef22a-abc12 container=adapter benchmark_id=arc_easy ===
INFO starting evaluation
INFO benchmark completed

Get Evaluation Benchmark Logs

Returns plain-text workload logs for a single benchmark within an evaluation job. The benchmark is identified by benchmark_index in the request path.

Kubernetes runtime: adapter container stdout/stderr via the Kubernetes API. Local runtime: contents of the benchmark's jobrun.log file. See GET /api/v1/evaluations/jobs/{id}/logs for shared query parameters.

path Parameters
id
required
string (Id)
benchmark_index
required
integer (Benchmark Index) >= 0
query Parameters
-1 (integer) or integer
Default: 1000

Maximum number of log lines to return. Use -1 to request all available log lines (no tail limit). Valid values: 1–10000, or -1 for all lines.

timestamps
boolean
Default: false

Include Kubernetes log timestamps

since_seconds
integer >= 1

Only return logs newer than this many seconds

Responses

Response samples

Content type
text/plain
INFO starting evaluation
INFO benchmark completed

Create Post-Processing Computation

Create an asynchronous post-processing computation.

Request Body schema: application/json
required
name
string (PostProcessingName) non-empty

Name of the confidence interval post-processing computation.

object (PostProcessingHardwareConfig)

Optional hardware override for the post-processing job.

required
Array of objects (StandalonePostProcessingOperations) non-empty

Post-processing operations executed by a standalone resource.

Responses

Request samples

Content type
application/json
{
  • "name": "inspect-accuracy-confidence-interval",
  • "hardware_config": {
    },
  • "operations": [
    ]
}

Response samples

Content type
application/json
{
  • "name": "string",
  • "hardware_config": {
    },
  • "resource": {
    },
  • "operations": [
    ],
  • "status": {
    },
  • "results": {
    }
}

Get Post-Processing Computation

Get a standalone confidence interval post-processing resource.

path Parameters
id
required
string

Responses

Response samples

Content type
application/json
Example
{
  • "resource": {
    },
  • "name": "inspect-accuracy-confidence-interval",
  • "hardware_config": {
    },
  • "operations": [
    ],
  • "status": {
    },
  • "results": {
    }
}

Delete Post-Processing Computation

Delete a standalone confidence interval post-processing resource.

path Parameters
id
required
string

Responses

Response samples

Content type
application/json
{
  • "message": "The bearer token is not valid.",
  • "message_code": "invalid_auth_token",
  • "trace": "b12692e1-8582-4628-88ca-7a13fefb73e2"
}

Create Confidence Interval Post-Processing

Create confidence interval computation for the benchmarks in an existing evaluation job. The operation is idempotent: if the same job already has a confidence interval computation, the existing post-processing resource is returned.

The computation runs independently for each benchmark. Each item in operations describes one post-processing operation. hardware_config applies to the post-processing job as a whole.

path Parameters
job_id
required
string (Job Id)
Request Body schema: application/json
required
name
string (PostProcessingName) non-empty

Name of the confidence interval post-processing computation.

object (PostProcessingHardwareConfig)

Optional hardware override for the post-processing job.

required
Array of objects (JobPostProcessingOperations) non-empty

Post-processing operations associated with an evaluation job.

Responses

Request samples

Content type
application/json
{
  • "name": "inspect-accuracy-confidence-interval",
  • "hardware_config": {
    },
  • "operations": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "inspect-accuracy-confidence-interval",
  • "hardware_config": {
    },
  • "operations": [
    ],
  • "status": {
    },
  • "results": {
    }
}

Get Job Post-Processing Computation

Get a confidence interval computation associated with an evaluation job.

path Parameters
job_id
required
string (Job Id)
post_processing_id
required
string (Post Processing Id)

Responses

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "inspect-accuracy-confidence-interval",
  • "hardware_config": {
    },
  • "operations": [
    ],
  • "status": {
    },
  • "results": {
    }
}

Delete Job Post-Processing Computation

Delete a confidence interval computation associated with an evaluation job.

path Parameters
job_id
required
string (Job Id)
post_processing_id
required
string (Post Processing Id)

Responses

Response samples

Content type
application/json
{
  • "message": "The bearer token is not valid.",
  • "message_code": "invalid_auth_token",
  • "trace": "b12692e1-8582-4628-88ca-7a13fefb73e2"
}

Collections

Benchmark collection management endpoints

List Collections

List all benchmark collections.

query Parameters
limit
integer (Limit) [ 1 .. 100 ]
Default: 50

Maximum number of collections to return

offset
integer (Offset) >= 0
Default: 0

Offset for pagination

name
string (Name)

Name to search for

category
string (Category)

Category to search for

tags
string (Tags)

Tags to search for

scope
string (Scope of collections)
Enum: "system" "tenant"

Filters the result set by collection scope, within the collections visible to the requesting tenant. system returns all system-defined collections (including curated ones, identified by curation_order > 0). tenant returns only this tenant's own custom collections. When omitted, all collections visible to this tenant are returned (system + this tenant's custom collections).

modalities
string (Modality filter)

Filter collections by modality (snake_case). Returns collections whose modalities array contains this value. May be repeated for multiple values.

tasks
string (Task filter)

Filter collections by task type (snake_case). Returns collections whose tasks array contains this value. May be repeated for multiple values.

domains
string (Domain filter)

Filter collections by evaluation domain (snake_case). Returns collections whose domains array contains this value. May be repeated for multiple values.

industries
string (Industry filter)

Filter collections by target industry (snake_case). Returns collections whose industries array contains this value. May be repeated for multiple values.

evaluation_targets
string (Evaluation target filter)

Filter collections by evaluation target type (snake_case). Example values: model, agent. Returns collections whose evaluation_targets array contains this value. May be repeated for multiple values.

sort_by
string (Sort order)
Value: "curation_order"

Sort field for the result set. Applied server-side across all matching collections before pagination. curation_order sorts curated collections ascending (0 or absent last). When omitted, the existing default collection ordering is preserved.

Responses

Response samples

Content type
application/json
{
  • "first": {
    },
  • "limit": 50,
  • "total_count": 2,
  • "items": [
    ]
}

Create Collection

Create a new collection.

Request Body schema: application/json
required
Any of
name
required
string

Collection name.

category
required
string [ 1 .. 128 ] characters
Deprecated

DEPRECATED — use domains instead. Retained for backwards compatibility. A create or update request must include either this field or a non-empty domains array.

description
string

Optional description.

tags
Array of strings

Tags.

object

Custom key-value data.

object (PassCriteria)

Pass criteria for the collection.

required
Array of objects (CollectionBenchmarkConfig)

Benchmarks in the collection.

curation_order
integer

Controls the listing priority of this collection within collection retrieval results. 0 (or absent) means not curated. Positive integers specify priority — lower values have higher priority. Set by system operators via YAML configuration only; the server rejects writes from tenant API consumers.

domains
Array of strings

High-level evaluation domains this collection addresses (snake_case). If not explicitly set, the handler returns the union of domains from the collection's benchmarks (via BenchmarkResource.domains). Canonical values: knowledge_and_reasoning, grounded_document_understanding, instruction_and_output_reliability, tool_use_and_function_calling, software, trustworthiness, multilingual, multimodal.

tasks
Array of strings

ML tasks this collection evaluates (snake_case). If not explicitly set, the handler returns the union of tasks from the collection's benchmarks (via BenchmarkResource.tasks). Known values: reasoning, data_analysis, extraction, summarization, full_document_qa, long_context_understanding, rag, citation_attribution, grounding_discipline, instruction_following, structured_output, constraint_following, call_generation, code_generation, code_understanding, code_repair, safety, calibration, abstention, adversarial_injection, translation, document_chart_vqa, general_visual_reasoning.

modalities
Array of strings

Data modalities this collection covers (snake_case). If not explicitly set, the handler returns the union of modalities from the collection's benchmarks (via BenchmarkResource.modalities). Known values: text, vision, multimodal.

industries
Array of strings

Business industries this collection is relevant for (snake_case). Collection-level field; the same benchmarks may serve different industries. Example values: health, telco, financial, government.

evaluation_targets
Array of strings

AI entity types this collection evaluates (snake_case). If not explicitly set, the handler returns the union of evaluation_targets from the collection's benchmarks (via BenchmarkResource.evaluation_targets). Example values: model, agent.

object (CollectionAgentMetadata)

Structured metadata for AI agent discoverability at the collection level.

Responses

Request samples

Content type
application/json
Example
{
  • "name": "release-gate-safety",
  • "category": "safety",
  • "description": "Release-gate collection combining reasoning and red-teaming benchmarks",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "release-gate-safety",
  • "category": "safety",
  • "description": "Release-gate collection combining reasoning and red-teaming benchmarks",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Get Collection

Get details of a specific collection.

path Parameters
id
required
string (Collection Id)

Responses

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "llm-safety-suite",
  • "category": "safety",
  • "description": "Comprehensive safety evaluation combining reasoning accuracy and vulnerability scanning",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Update Collection

Update an existing collection.

path Parameters
id
required
string (Collection Id)
Request Body schema: application/json
required
Any of
name
required
string

Collection name.

category
required
string [ 1 .. 128 ] characters
Deprecated

DEPRECATED — use domains instead. Retained for backwards compatibility. A create or update request must include either this field or a non-empty domains array.

description
string

Optional description.

tags
Array of strings

Tags.

object

Custom key-value data.

object (PassCriteria)

Pass criteria for the collection.

required
Array of objects (CollectionBenchmarkConfig)

Benchmarks in the collection.

curation_order
integer

Controls the listing priority of this collection within collection retrieval results. 0 (or absent) means not curated. Positive integers specify priority — lower values have higher priority. Set by system operators via YAML configuration only; the server rejects writes from tenant API consumers.

domains
Array of strings

High-level evaluation domains this collection addresses (snake_case). If not explicitly set, the handler returns the union of domains from the collection's benchmarks (via BenchmarkResource.domains). Canonical values: knowledge_and_reasoning, grounded_document_understanding, instruction_and_output_reliability, tool_use_and_function_calling, software, trustworthiness, multilingual, multimodal.

tasks
Array of strings

ML tasks this collection evaluates (snake_case). If not explicitly set, the handler returns the union of tasks from the collection's benchmarks (via BenchmarkResource.tasks). Known values: reasoning, data_analysis, extraction, summarization, full_document_qa, long_context_understanding, rag, citation_attribution, grounding_discipline, instruction_following, structured_output, constraint_following, call_generation, code_generation, code_understanding, code_repair, safety, calibration, abstention, adversarial_injection, translation, document_chart_vqa, general_visual_reasoning.

modalities
Array of strings

Data modalities this collection covers (snake_case). If not explicitly set, the handler returns the union of modalities from the collection's benchmarks (via BenchmarkResource.modalities). Known values: text, vision, multimodal.

industries
Array of strings

Business industries this collection is relevant for (snake_case). Collection-level field; the same benchmarks may serve different industries. Example values: health, telco, financial, government.

evaluation_targets
Array of strings

AI entity types this collection evaluates (snake_case). If not explicitly set, the handler returns the union of evaluation_targets from the collection's benchmarks (via BenchmarkResource.evaluation_targets). Example values: model, agent.

object (CollectionAgentMetadata)

Structured metadata for AI agent discoverability at the collection level.

Responses

Request samples

Content type
application/json
Example
{
  • "name": "llm-safety-suite",
  • "category": "safety",
  • "description": "Safety evaluation with reasoning, OWASP risks, and content quality",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "llm-safety-suite",
  • "category": "safety",
  • "description": "Safety evaluation with reasoning, OWASP risks, and content quality",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Patch Collection

Partially update an existing collection.

path Parameters
id
required
string (Collection Id)
Request Body schema: application/json
required
Array
op
required
string (PatchOp)
Enum: "replace" "add" "remove"

Patch operation type

path
required
string

JSON Pointer path

value
any

Value for add/replace (omit for remove)

Responses

Request samples

Content type
application/json
Example
[
  • {
    },
  • {
    }
]

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "llm-safety-suite",
  • "category": "safety",
  • "description": "Safety evaluation with stricter pass threshold",
  • "tags": [
    ],
  • "pass_criteria": {
    },
  • "benchmarks": [
    ]
}

Delete Collection

Delete a collection.

path Parameters
id
required
string (Collection Id)

Responses

Response samples

Content type
application/json
{
  • "message": "The field 'state' is not valid.",
  • "message_code": "invalid_value",
  • "trace": "b12692e1-8582-4628-88ca-7a13fefb73e2"
}

Copy Collection

Creates a new tenant-scoped (custom) collection as a copy of the specified source collection. The server copies all CollectionConfig fields from the source, assigns a new ID with scope=tenant, and records the source ID in derived_from for provenance. The request body may include optional overrides (e.g. a new name).

path Parameters
id
required
string (Collection Id)

ID of the source collection to copy.

Request Body schema: application/json
optional

Optional field overrides for the cloned collection. Any field omitted here is inherited from the source collection, except curation_order. Clone requests never accept curation_order, and the server always resets it to 0.

name
string

New name for the copy. Defaults to source name.

description
string

Override description.

category
string

Deprecated — use domains.

tags
Array of strings

Override tags.

object (PassCriteria)

Override pass criteria.

Array of objects (CollectionBenchmarkConfig)

Override benchmarks.

domains
Array of strings

Override domains.

tasks
Array of strings

Override tasks.

modalities
Array of strings

Override modalities.

industries
Array of strings

Override industries.

evaluation_targets
Array of strings

Override evaluation targets.

object

Override custom key-value data.

object (CollectionAgentMetadata)

Override agent metadata for AI agent consumption.

Responses

Request samples

Content type
application/json
{
  • "name": "my-custom-rag-eval",
  • "domains": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "derived_from": "e5f6a7b8-9012-3456-cdef-0123456789ab",
  • "pinned_order": 0,
  • "status": {
    },
  • "name": "my-rag-evaluation",
  • "category": "document_understanding",
  • "tasks": [
    ],
  • "modalities": [
    ],
  • "benchmarks": [
    ]
}

Providers

Evaluation provider endpoints

List Providers

List all registered evaluation providers.

query Parameters
limit
integer (Limit) [ 1 .. 100 ]
Default: 50

Maximum number of providers to return

offset
integer (Offset) >= 0
Default: 0

Offset for pagination

benchmarks
boolean (Benchmarks)
Default: true

Include or exclude benchmarks supported by this provider in the response

name
string (Name)

Name to search for

tags
string (Tags)

Tags to search for

scope
string (Scope of providers)
Enum: "system" "tenant"

Set to system to get only system defined providers, or tenant to get only user defined providers. If scope is not provided, both system and user defined providers will be returned.

Responses

Response samples

Content type
application/json
{
  • "first": {
    },
  • "limit": 50,
  • "total_count": 3,
  • "items": [
    ]
}

Create a new provider scoped to the current tenant (Bring Your Own Provider)

Create a new provider scoped to the current tenant (Bring Your Own Provider)

Request Body schema: application/json
required
name
required
string

Provider name

title
string

Provider display title

description
string

Provider description

tags
Array of strings

Provider tags

object (AgentMetadata)

Agent discoverability metadata for this provider

required
object (Runtime)

Provider runtime configuration

required
Array of objects (BenchmarkResource)

Benchmarks offered by this provider

Responses

Request samples

Content type
application/json
Example
{
  • "name": "my-custom-evaluator",
  • "title": "Custom Internal Evaluator",
  • "description": "Internal evaluation adapter for domain-specific benchmarks",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "my-custom-evaluator",
  • "title": "Custom Internal Evaluator",
  • "description": "Internal evaluation adapter for domain-specific benchmarks",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Get Provider

Get a provider by ID.

path Parameters
id
required
string (Provider Id)

Provider ID

Responses

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "garak",
  • "title": "Garak",
  • "description": "LLM vulnerability scanner and red-teaming framework",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Update Provider

Update an existing provider.

path Parameters
id
required
string (Provider Id)

Provider ID

Request Body schema: application/json
required
name
required
string

Provider name

title
string

Provider display title

description
string

Provider description

tags
Array of strings

Provider tags

object (AgentMetadata)

Agent discoverability metadata for this provider

required
object (Runtime)

Provider runtime configuration

required
Array of objects (BenchmarkResource)

Benchmarks offered by this provider

Responses

Request samples

Content type
application/json
{
  • "name": "my-custom-evaluator",
  • "title": "Custom Internal Evaluator",
  • "description": "Updated evaluation adapter with improved tokenization",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "my-custom-evaluator",
  • "title": "Custom Internal Evaluator",
  • "description": "Updated evaluation adapter with improved tokenization",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Patch Provider

Partially update an existing provider.

path Parameters
id
required
string (Provider Id)
Request Body schema: application/json
required
Array
op
required
string (PatchOp)
Enum: "replace" "add" "remove"

Patch operation type

path
required
string

JSON Pointer path

value
any

Value for add/replace (omit for remove)

Responses

Request samples

Content type
application/json
Example
[
  • {
    },
  • {
    }
]

Response samples

Content type
application/json
{
  • "resource": {
    },
  • "name": "my-custom-evaluator",
  • "title": "Custom Internal Evaluator",
  • "description": "Updated evaluation adapter with bug fixes",
  • "tags": [
    ],
  • "runtime": {
    },
  • "benchmarks": [
    ]
}

Delete Provider

Delete provider by ID.

path Parameters
id
required
string (Provider Id)

Provider ID

Responses

Response samples

Content type
application/json
{
  • "message": "The field 'state' is not valid.",
  • "message_code": "invalid_value",
  • "trace": "b12692e1-8582-4628-88ca-7a13fefb73e2"
}

Health

Health check endpoints

Health Check

Health check endpoint suitable for liveness and readiness probes.

Responses

Response samples

Content type
application/json
{
  • "status": "healthy",
  • "timestamp": "2026-05-27T18:42:11Z"
}