RAG Eval Sidekick evaluates outputs from retrieval-augmented generation (RAG) pipelines without requiring a person to inspect every answer manually. Give it one or more triples of:
{question, retrieved_chunks, answer}
It scores each triple for faithfulness, answer relevance, and context precision, then explains likely root causes such as irrelevant retrieval or an unsupported claim in the generated answer. It also includes retrieval configuration tuning and SQLite-backed regression history.
The product has a Streamlit interface at http://localhost:8501 and a FastAPI API at http://localhost:8000.
Demo video: https://www.youtube.com/watch?v=ZG2_PnxQcJs
Prerequisites: Docker Desktop, or Docker Engine with Docker Compose.
cp .env.example .envOpen .env and replace the placeholder with your OpenAI API key:
OPENAI_API_KEY=sk-your-openai-api-key-hereBuild and start the backend and frontend:
docker compose up --buildOlder installations may use docker-compose up --build instead. When both
services are ready, open:
- Streamlit application: http://localhost:8501
- FastAPI documentation: http://localhost:8000/docs
- FastAPI health check: http://localhost:8000/health
Stop the services with Ctrl+C, then run:
docker compose downRun history is stored in the named Docker volume rag_eval_history, so normal
container recreation does not erase saved runs.
The repository includes five prepared examples in
sample_data/example_triples.json:
- Two clean examples that fully pass the evaluation.
- One example with an organic, unplanned faithfulness failure: the retrieved chunks do not actually support the answer's specific claim about the definitions of BERT's masked language modeling (MLM) and next sentence prediction (NSP) tasks.
- One deliberately engineered retrieval failure with an irrelevant chunk swapped into the retrieved context.
- One deliberately engineered faithfulness failure with a fabricated statistic injected into the answer.
After starting the application:
- Open the 🔎 Evaluate tab.
- Click Load example triples.
- Click Run evaluation.
- Review the aggregate pass/fail summary, per-metric scores, diagnoses, and retrieved chunks.
The examples use the included Wikipedia-derived corpus at
sample_data/source_text.txt.
The two input paths serve different purposes and should not be confused.
Use Custom triple, POST /score-and-diagnose, or POST /report when you
already have RAG outputs. Sidekick does not require your application to use a
particular vector database, embedding model, framework, chunking strategy, or
generation model. Export each real result as:
{
"question": "What did the policy change?",
"retrieved_chunks": ["First retrieved passage", "Second retrieved passage"],
"answer": "The answer produced by your RAG system"
}This is the actual product workflow: Sidekick evaluates the behavior of an existing RAG system independently of how that system was built.
The Source document uploader accepts a UTF-8 .txt file and a question. It
runs Sidekick's small built-in RAG pipeline to chunk the document, embed and
retrieve passages, and generate an answer before evaluating the result. This path
is useful for demos, onboarding, and producing test data when no external RAG
pipeline is available. It is not required to evaluate your own system.
After uploading a document, the Suggest questions button uses GPT-5.6 to
propose four evaluation questions—a mix of straightforward factual questions and
at least one that requires connecting information across different parts of the
document. Clicking a suggestion automatically fills the question field.
For every RAG triple, Sidekick returns three scores from 0.0 to 1.0:
- Faithfulness: whether factual claims in the answer are supported by the retrieved chunks. GPT-5.6 applies an explicit rubric; a specific unsupported factual or numerical claim caps the score even if the rest of the answer is well-supported.
- Answer relevance: whether the answer directly and completely addresses the question, judged with a separate GPT-5.6 rubric.
- Context precision: the average embedding cosine similarity between the question and retrieved chunks, used as a practical estimate of retrieval relevance.
A score below 0.6 fails. Failed triples receive a short diagnosis distinguishing
retrieval problems, unsupported generation, and off-topic or incomplete answers.
Batch reports include pass/fail counts, metric averages, and non-exclusive failure
type counts.
Relevant endpoints:
POST /score-and-diagnosePOST /reportPOST /generate-triplefor the optional built-in demo pipelinePOST /suggest-questionsfor questions grounded in an uploaded source document
The ⚙️ Auto-tune tab evaluates combinations of chunk size and retrieval count
(top_k) across a set of questions. Defaults are:
chunk_sizes = [100, 150, 200, 300]
top_ks = [2, 3, 5]
Smaller custom lists can be supplied for faster tests. The tuner averages all
three evaluation metrics with equal weight, highlights the best configuration,
and explains how much better its combined score is than the worst configuration.
It batches all question embeddings once and reuses chunk embeddings across
top_k values to avoid redundant embedding calls.
Relevant endpoint: POST /tune.
The 🕘 Run History tab saves aggregate reports with unique labels such as
v1-baseline or v2-smaller-chunks. Saved runs include a UTC timestamp and the
full report. Two labeled runs can be compared side by side; score differences are
reported as run B minus run A, so positive values indicate improvement in the
second run.
History is stored locally in SQLite (runs.db outside Docker, or the persistent
Compose volume inside Docker).
Relevant endpoints:
POST /save-runGET /runsGET /compare-runs?label_a=...&label_b=...
Requires Python 3.11 or newer.
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .envIn manual mode, the application does not automatically load .env. Export the
key into both service terminals or their shared parent shell:
export OPENAI_API_KEY="your-key"Start the backend:
uvicorn backend.app:app --reloadIn a second terminal with the virtual environment active, start Streamlit:
BACKEND_URL=http://localhost:8000 streamlit run frontend/app.pyOpen http://localhost:8501. The API documentation is at http://localhost:8000/docs.
GPT-5.6 is the runtime model that powers the product's generation, evaluation, and diagnosis logic. This is distinct from Codex's role as the coding agent used to build the implementation.
- Faithfulness and answer-relevance judging: For each RAG triple, GPT-5.6
evaluates the answer using the explicit rubrics in
backend/scorer.py. The faithfulness judge compares claims against the retrieved chunks, while the answer-relevance judge evaluates how directly the answer addresses the question. Structured outputs constrain both scores to the0.0–1.0range. - Root-cause diagnosis: When any score is below
0.6, GPT-5.6 uses the prompt inbackend/diagnoser.pyto produce a short plain-English explanation. It distinguishes likely retrieval failures from unsupported answer claims and answer-relevance failures. - Built-in mini RAG generation: GPT-5.6 generates the answer in
backend/mini_rag.pyafter the built-in pipeline chunks, embeds, and retrieves source passages. This is the generation path used for the Source document demo workflow and for producing the prepared sample triples. - Auto-tuning generations: During configuration sweeps,
backend/tuner.pyinvokes the same GPT-5.6 answer-generation step for every question and each chunk-size/top_kcombination. Those generated answers are then scored so configurations can be compared on their resulting RAG quality rather than retrieval similarity alone. - Question suggestion: When generating a triple from an uploaded document,
GPT-5.6 reads the full source text and proposes four evaluation questions via the
/suggest-questionsendpoint inbackend/app.py, using structured output to guarantee exactly four unique, non-blank questions.
This project was built iteratively with Codex as an implementation and testing partner, rather than generated in one pass.
Codex accelerated:
- Scaffolding the FastAPI, Streamlit, sample-data, Docker, and Compose structure.
- Implementing the Wikipedia API downloader and the self-contained mini RAG pipeline, including custom document support.
- Turning the evaluation criteria into explicit, inspectable LLM-judge prompts and constrained structured outputs.
- Wiring scoring, diagnosis, aggregate reporting, multipart upload, tuning, SQLite history, and the tabbed frontend through typed API contracts.
- Optimizing the tuner to batch question embeddings and reuse chunk embeddings.
- Running offline fake-client tests and endpoint integration tests throughout. This caught issues during development such as missing dependencies, multipart request validation, API request-shape mismatches, and a floating-point equality mistake in a recommendation test. Codex also repeatedly boot-tested the Streamlit app and validated the Docker Compose configuration.
The product and evaluation decisions remained human-directed:
- Selecting faithfulness, answer relevance, and context precision as the core metrics.
- Defining the
0.6health threshold and deciding that diagnoses should separate retrieval failures from generation failures. - Reviewing the judge's behavior and identifying that the first faithfulness rubric was too lenient toward a confident fabricated statistic. The rubric was then tightened so unsupported falsifiable claims cannot receive the same score as a harmless nuance.
- Choosing the chunk-size and
top_ksweep ranges. - Expanding the initial evaluator into an auto-tuner and labeled regression-history tool.
- Keeping the built-in mini RAG path explicitly separate from the primary use case of evaluating outputs from any external RAG pipeline.
The project code is released under the MIT License. See LICENSE.
The included source corpus contains text from the following English Wikipedia articles:
That source material is used under the
Creative Commons Attribution-ShareAlike 4.0 International license.
The attribution is also included at the end of
sample_data/source_text.txt. The Wikipedia-derived
content remains subject to CC BY-SA 4.0; the project's MIT license applies to the
project code and does not replace the source material's license.