Evaluating LLM Components in RAG Applications with DeepEval
This post demonstrates how to use DeepEval to systematically evaluate different components of a Retrieval-Augmented Generation (RAG) application, including retrievers, generators, and the full pipeline. It covers preparing golden datasets, applying relevant metrics like recall, precision, and faithfulness, and leveraging G-Eval for holistic application assessment.
Evaluating a RAG (Retrieval-Augmented Generation) application is not just about checking the final answer. Each component—retriever, reranker, generator—can fail in its own unique way. DeepEval gives me a way to isolate and measure those failures, using targeted evaluation metrics and carefully curated 'golden' datasets (reference answers). Here's how I do it.
Why Component-Level Evaluation Matters
When a RAG system gives a bad answer, is it because:
The retriever fetched irrelevant docs?
The reranker picked the wrong ones?
The generator hallucinated?
If I only look at the final output, I can't tell. DeepEval lets me test each part in isolation, using targeted metrics:
Component
Key Metrics
What It Measures
Retriever
Contextual Recall, Contextual Precision
Did it fetch the right context?
Generator
Faithfulness, Answer Relevancy
Did it hallucinate? Was it on-topic?
Pipeline
Correctness, Completeness, Style (G-Eval)
Is the final answer accurate, full, clear?
Prepping Golden Datasets with DeepEval Synthesizer
The first step is to build a "golden" dataset: queries, ideal answers, and (sometimes) the exact context that should be retrieved. DeepEval's synthesizer can generate these, but in my experience, they often need review and refinement. I always:
Review and edit the generated goldens for clarity and coverage.
Remove ambiguous or trivial questions.
Add extra context or edge cases that matter for my use case.
This curation step is critical. Your evaluation results will only be as reliable as the quality and coverage of your golden set.
Tip: I keep my goldens in JSON, versioned in git. Every time I tweak the retriever or generator, I rerun the evals to see real progress (or regressions).
Evaluating the Retriever
To test the retriever, I check: when given a query, does it fetch the chunks that actually contain the answer? I use DeepEval's ContextualRecallMetric and ContextualPrecisionMetric for this.
# eval_retriever.pyimport json
from dotenv import load_dotenv
from deepeval import evaluate
from deepeval.test_case import LLMTestCase
from deepeval.metrics import ContextualRecallMetric, ContextualPrecisionMetric
from src.retriever import build_retriever
load_dotenv()
GOLDEN_PATH = "goldens/retriever_deepeval_goldens.json"
JUDGE_MODEL = "gpt-4o-mini"
THRESHOLD = 0.7defload_goldens(path):
withopen(path, encoding="utf-8") as f:
return json.load(f)
defmain():
retriever = build_retriever()
goldens = load_goldens(GOLDEN_PATH)
test_cases = []
for g in goldens:
retrieved = retriever.invoke(g["query"])
retrieval_context = [doc.page_content for doc in retrieved]
test_cases.append(
LLMTestCase(
input=g["query"],
expected_output=g["ideal_answer"],
retrieval_context=retrieval_context,
actual_output="(generator not evaluated in this run)",
)
)
metrics = [
ContextualRecallMetric(threshold=THRESHOLD, model=JUDGE_MODEL, include_reason=True),
ContextualPrecisionMetric(threshold=THRESHOLD, model=JUDGE_MODEL, include_reason=True),
]
evaluate(test_cases=test_cases, metrics=metrics, hyperparameters={
"retriever": "base_k5",
"embedding_model": "text-embedding-3-small",
"chunk_size": 1000,
"chunk_overlap": 150,
"top_k": 5,
"judge_model": JUDGE_MODEL,
"golden_set": GOLDEN_PATH,
},)
if __name__ == "__main__":
main()
I can swap in a reranker, change top-k, or try new embeddings, and quickly see if recall/precision improve.
Evaluating the Generator
For the generator, I want to know: given the right context, does it stick to the facts, or does it make things up? I feed it the golden context and use FaithfulnessMetric and AnswerRelevancyMetric.
"""
evals/eval_generator.py
=======================
Component-level evaluation of the GENERATOR, in isolation.
Faithfulness: of the claims in the generated answer, how many are supported
by the context it was given? (Did the generator make things up?)
ISOLATION: we feed the generator the GOLDEN context (the known-good chunks
from the faithfulness dataset), NOT the retriever's output. So a low score
is purely the generator's fault --- the context was already correct.
python -m evals.eval_generator
"""import json
from dotenv import load_dotenv
from deepeval import evaluate
from deepeval.test_case import LLMTestCase
from deepeval.metrics import FaithfulnessMetric, AnswerRelevancyMetric
from src.generator import generate # your generator: generate(query, context) -> answer
load_dotenv()
GOLDEN_PATH = "goldens/faithfulness_dataset.json"
JUDGE_MODEL = "gpt-4o-mini"
THRESHOLD = 0.7defload_goldens(path):
withopen(path, encoding="utf-8") as f:
return json.load(f)
defmain():
goldens = load_goldens(GOLDEN_PATH)
test_cases = []
for g in goldens:
context = [g["context"]]
answer = generate(g["query"], context)
test_cases.append(
LLMTestCase(
input=g["query"],
actual_output=answer,
retrieval_context=context,
)
)
metrics = [
FaithfulnessMetric(
threshold=THRESHOLD,
model=JUDGE_MODEL,
include_reason=True,
),
AnswerRelevancyMetric(
threshold=THRESHOLD,
model=JUDGE_MODEL,
include_reason=True,
),
]
evaluate(test_cases=test_cases, metrics=metrics)
if __name__ == "__main__":
main()
If faithfulness drops in this setup, it's likely due to the generator, since we're providing the correct context.
Evaluating the Whole Pipeline with G-Eval
Finally, I want to know: does the whole system answer questions correctly, completely, and in the right tone? For this, I use DeepEval's G-Eval metrics (a set of LLM-based evaluation rubrics): Correctness, Completeness, and Style.
# eval_application.pyimport json
from dotenv import load_dotenv
from deepeval import evaluate
from deepeval.test_case import LLMTestCase, LLMTestCaseParams
from deepeval.metrics import GEval
from deepeval.metrics.g_eval import Rubric
from src.rag_pipeline import RagPipeline
load_dotenv()
GOLDEN_PATH = "goldens/faithfulness_dataset.json"
JUDGE_MODEL = "gpt-4o-mini"
THRESHOLD = 0.7defload_goldens(path):
withopen(path, encoding="utf-8") as f:
return json.load(f)
defmain():
rag = RagPipeline()
goldens = load_goldens(GOLDEN_PATH)
test_cases = []
for g in goldens:
result = rag.invoke(g["query"])
test_cases.append(
LLMTestCase(
input=g["query"],
actual_output=result["answer"],
expected_output=g["ideal_answer"],
)
)
correctness = GEval(
name="Correctness",
evaluation_steps=[
"Compare only the factual claims in the actual output against the expected output.",
"A claim is wrong only if it CONTRADICTS the expected output or is factually false. Judge truth, not completeness.",
"A factually accurate answer must score at least 0.9 even if it is shorter or covers fewer points than the expected output.",
"Do NOT deduct for brevity, missing elaboration, or omitted points --- omissions are not errors here.",
"Additional correct information must NEVER lower the score.",
],
rubric=[
Rubric(score_range=(9, 10), expected_outcome="All stated claims are factually correct and consistent. No contradictions. Brevity is fine."),
Rubric(score_range=(5, 8), expected_outcome="Mostly correct but one minor inaccuracy."),
Rubric(score_range=(0, 4), expected_outcome="Contains a clear factual error or a claim that contradicts the expected output."),
],
evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
threshold=THRESHOLD,
model=JUDGE_MODEL,
strict_mode=False,
)
completeness = GEval(
name="Completeness",
evaluation_steps=[
"Identify the key points contained in the expected output.",
"Check how many of those key points are addressed in the actual output.",
"Penalize the actual output for each key point from the expected output that it omits or only partially covers.",
"Judge coverage only. Do NOT lower the score because a covered point is stated incorrectly --- factual correctness is judged separately.",
"Do NOT penalize the actual output for adding extra information beyond the expected output.",
],
rubric=[
Rubric(score_range=(9, 10), expected_outcome="Addresses essentially all key points in the expected output."),
Rubric(score_range=(5, 8), expected_outcome="Covers the main key points but misses one or more."),
Rubric(score_range=(0, 4), expected_outcome="Misses several key points; only partially covers the expected output."),
],
evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT, LLMTestCaseParams.EXPECTED_OUTPUT],
threshold=THRESHOLD,
model=JUDGE_MODEL,
strict_mode=False,
)
style = GEval(
name="Style",
evaluation_steps=[
"Judge only the teaching style and tone of the actual output, not whether it is factually correct or complete.",
"Reward an intuitive, explanatory tone: plain language, the idea explained before any formula or jargon, and technical terms briefly unpacked when used.",
"Reward a direct, conversational register written in prose, as an experienced AI engineer would explain it out loud, rather than a dry, formal, or bullet-list tone.",
"An analogy or concrete example is a BONUS when the concept is abstract, but a clear, direct, well-explained answer is fully acceptable and must NOT be penalized for not having one.",
"Penalize answers that are stiff, bureaucratic, structured as a bare list with no explanation, or that use unexplained jargon.",
"Do NOT reward or penalize based on correctness, completeness, or length --- only on style and tone.",
],
rubric=[
Rubric(score_range=(9, 10), expected_outcome="Clearly in an experienced AI engineer's teaching voice: intuitive, conversational prose that explains before it formalizes."),
Rubric(score_range=(7, 8), expected_outcome="Clear, conversational, and well-explained in prose. Fully acceptable even without an analogy or example."),
Rubric(score_range=(4, 6), expected_outcome="Understandable but somewhat flat, formal, or list-heavy in places."),
Rubric(score_range=(0, 3), expected_outcome="Dry, stiff, bare-list, jargon-heavy, or robotic; does not read like a teaching explanation."),
],
evaluation_params=[LLMTestCaseParams.INPUT, LLMTestCaseParams.ACTUAL_OUTPUT],
threshold=THRESHOLD,
model=JUDGE_MODEL,
strict_mode=False,
)
evaluate(test_cases=test_cases, metrics=[correctness, completeness, style])
if __name__ == "__main__":
main()
This is where I see the real user-facing quality. If correctness is high but completeness is low, maybe my retriever is missing key docs. If the style feels robotic or unnatural, I tweak the prompt to encourage a more conversational tone.
My Takeaways
Golden datasets are everything. Spend time curating them; don't trust any auto-generated set blindly.
Isolate before you optimize. Always measure retriever and generator separately before tuning the pipeline.
Metrics reveal the real pain points. If faithfulness is high but recall is low, it's usually the retriever; if recall is high but correctness is low, it's often the generator.
Automate and repeat. I rerun these evals on every major change. It's the only way to know if I'm actually making progress.
DeepEval isn't magic, but it's the most effective tool I've used so far for real, actionable RAG evaluation. If you're serious about quality, start with goldens and measure everything.
Join the discussion
Nothing here yet — be the first to weigh in.