Skip to content

Benchmarks Reference

The DeepEval adapter provides 8 benchmarks across two evaluation modes. Single-turn benchmarks evaluate individual input–output pairs. Multi-turn benchmarks evaluate sequences of conversation turns represented as ConversationalTestCase objects.

Single-turn benchmarks load dataset rows as LLMTestCase objects and apply a single DeepEval metric per run.

Benchmark IDNameCategoryPrimary Metric
faithfulnessFaithfulnessRAG evaluationfaithfulness_score
relevancyAnswer RelevancyRAG evaluationrelevancy_score
hallucinationHallucinationSafetyhallucination_score
correctnessCorrectnessAccuracycorrectness_score
summarizationSummarizationNLPsummarization_score

Tests whether the LLM output is supported by the provided retrieval context. Claims in the output are verified against the context; unsupported claims lower the score.

Required dataset columns:

ColumnTypeDescription
inputstringThe user query or prompt
actual_outputstringThe LLM-generated response
retrieval_contextstring (or JSON array)The retrieved context passages used to generate the response

Output metrics: faithfulness_score, claims_count, supported_claims_count

Use case: Verify that a RAG pipeline does not introduce claims beyond what the retrieved documents support.


Tests whether the LLM output is relevant to and directly addresses the input query.

Required dataset columns:

ColumnTypeDescription
inputstringThe user query or prompt
actual_outputstringThe LLM-generated response

Output metrics: relevancy_score

Use case: Detect responses that are grammatically correct but do not actually answer the question.


Tests for hallucinated content in the LLM output — statements that are not grounded in the provided context.

Required dataset columns:

ColumnTypeDescription
inputstringThe user query or prompt
actual_outputstringThe LLM-generated response
contextstring (or JSON array)Ground-truth context the response should be based on

Output metrics: hallucination_score, hallucination_detected

hallucination_detected is a boolean derived from whether hallucination_score exceeds the configured threshold.

Use case: Safety evaluation to detect when a model generates factually incorrect information not present in the source documents.


Tests factual correctness of the LLM output against a ground-truth reference answer. Uses GEval with a correctness criterion.

Required dataset columns:

ColumnTypeDescription
inputstringThe user query or prompt
actual_outputstringThe LLM-generated response
expected_outputstringThe ground-truth reference answer

Output metrics: correctness_score

Use case: Evaluate accuracy in question-answering tasks where ground-truth answers are available.


Tests the quality of LLM-generated summaries — whether the summary covers the key information from the source and is not hallucinated.

Required dataset columns:

ColumnTypeDescription
inputstringThe source text to be summarised
actual_outputstringThe LLM-generated summary

Output metrics: summarization_score

Use case: Evaluate summarisation pipelines for coverage of key information and alignment with the source text.


Multi-turn benchmarks evaluate full conversation sequences as ConversationalTestCase objects. Each row in the dataset represents one complete conversation, encoded as a list of turns.

Benchmark IDNameCategoryPrimary Metric
conversation-completenessConversation CompletenessMulti-turnconversation_completeness_score
role-adherenceRole AdherenceMulti-turnrole_adherence_score
knowledge-retentionKnowledge RetentionMulti-turnknowledge_retention_score

Tests whether the chatbot adequately addresses all user needs and tasks across the full conversation. The judge evaluates each turn’s handling of the user’s stated or implied intent.

Required dataset columns:

ColumnTypeDescription
turnsarray of turn objectsThe full conversation as a list of {"role": ..., "content": ...} objects

Optional dataset columns: chatbot_role, scenario, expected_outcome

Output metrics: conversation_completeness_score

Use case: Evaluate customer support or task-completion chatbots to ensure all user requests are handled rather than deflected or ignored.


Tests whether the chatbot stays in its assigned persona or role throughout the conversation. Requires the chatbot role to be specified either in the dataset or via the chatbot_role parameter.

Required dataset columns:

ColumnTypeDescription
turnsarray of turn objectsThe full conversation
chatbot_rolestringThe persona the chatbot should maintain (e.g. "helpful customer support agent") — required unless set via parameters.chatbot_role

Optional dataset columns: scenario

Output metrics: role_adherence_score

Use case: Evaluate persona-constrained chatbots (customer support, educational tutors, domain experts) to ensure the model does not break character.


Tests whether the chatbot correctly uses and retains information that the user disclosed in earlier turns of the conversation.

Required dataset columns:

ColumnTypeDescription
turnsarray of turn objectsThe full conversation

Optional dataset columns: chatbot_role, scenario

Output metrics: knowledge_retention_score

Use case: Evaluate stateful chatbot systems to confirm that user-provided facts (name, preferences, prior context) are remembered and applied correctly in later responses.


CSV is the default format (dataset_format: "csv") and is recommended for single-turn benchmarks. Each row is one test case.

Minimal faithfulness example:

input,actual_output,retrieval_context
"What is the boiling point of water?","Water boils at 100°C at sea level.","Water boils at 100 degrees Celsius (212°F) at standard atmospheric pressure."
"Who invented the telephone?","Alexander Graham Bell invented the telephone in 1876.","Alexander Graham Bell is credited with inventing the first practical telephone, patented in 1876."

JSONL (one JSON object per line) is the recommended format for multi-turn benchmarks because the turns field is represented as a native JSON array.

Basic turns (conversation-completeness, knowledge-retention):

{"turns": [{"role": "user", "content": "How do I reset my password?"}, {"role": "assistant", "content": "Click 'Forgot password' on the login page."}]}
{"turns": [{"role": "user", "content": "My account is locked."}, {"role": "assistant", "content": "I can help you unlock it. Can you provide your email address?"}]}

With chatbot_role (required for role-adherence):

{"turns": [{"role": "user", "content": "Hello"}, {"role": "assistant", "content": "Hi! How can I help you today?"}], "chatbot_role": "friendly customer support agent"}
{"turns": [{"role": "user", "content": "I want to cancel my subscription."}, {"role": "assistant", "content": "I understand. I can help you with that."}], "chatbot_role": "friendly customer support agent"}

With optional metadata:

{"turns": [{"role": "user", "content": "My name is Alice."}, {"role": "assistant", "content": "Nice to meet you, Alice!"}, {"role": "user", "content": "What is my name?"}, {"role": "assistant", "content": "Your name is Alice."}], "chatbot_role": "helpful assistant", "scenario": "User introduces themselves and expects the chatbot to remember"}

JSON format is equivalent to JSONL but stores all records as a top-level array. Use when your tooling produces JSON rather than JSONL.

[
{
"turns": [
{"role": "user", "content": "What is the capital of France?"},
{"role": "assistant", "content": "The capital of France is Paris."}
]
}
]
ColumnTypeRequired ByDescription
inputstringfaithfulness, relevancy, hallucination, correctness, summarizationUser query or prompt
actual_outputstringAll single-turn benchmarksLLM-generated response
retrieval_contextstring or arrayfaithfulnessRetrieved context passages
contextstring or arrayhallucinationGround-truth context to compare against
expected_outputstringcorrectnessGround-truth reference answer
turnsarray of {role, content}All multi-turn benchmarksFull conversation turn sequence
chatbot_rolestringrole-adherence (required), others (optional)Chatbot persona description
scenariostringOptional for all multi-turnHuman-readable description of the conversation scenario
expected_outcomestringOptional for conversation-completenessExpected outcome for completeness evaluation