rag-for-mdn-reference.vercel.app The demo requires an AI provider API key (e.g., Groq, OpenAI, Anthropic) and a Voyage API key for reranking. Both offer free tiers. No local setup needed — but you will need to configure keys in the Settings panel before asking questions.
This project answers a common question in AI engineering: why isn't plain vector search enough?
It builds a Retrieval-Augmented Generation (RAG) chatbot over the MDN JavaScript Guide using a 3-stage pipeline — hybrid vector + BM25 search, rank fusion, and neural reranking — to deliver cited, context-grounded answers to JavaScript questions.
This is a course project and proof of concept for production RAG patterns — not a hardened production system.
Semantic search over MDN docs with cited sources — the 3-stage pipeline in action.
CleanShot.2026-07-03.at.09.50.53-converted.mp4
- Live Demo (Bring Your Own Keys)
- What This Is
- Prerequisites
- Setup
- Database
- Development
- Available Scripts
- Tech Stack
- Acknowledgments
- Learn More
- Known Issues
- Rate Limits
bun iOption A: Docker Compose (cross-platform, recommended):
# Start PostgreSQL container
docker compose up -d
# Create the database and enable pgvector
docker compose exec postgres psql -U postgres -c "CREATE DATABASE \"unlearn-rag-course\";"
docker compose exec postgres psql -U postgres -d unlearn-rag-course -c "CREATE EXTENSION IF NOT EXISTS vector;"Option B: Local installation:
# macOS with Homebrew
brew install postgresql@18
brew install pgvector
# Start PostgreSQL
brew services start postgresql@18
# Create the database
createdb unlearn-rag-course
# Enable the pgvector extension
psql -d unlearn-rag-course -c "CREATE EXTENSION IF NOT EXISTS vector;"# Ubuntu/Debian
sudo apt install postgresql-18 postgresql-18-pgvector
# Start PostgreSQL
sudo systemctl start postgresql
# Create the database and enable pgvector
sudo -u postgres createdb unlearn-rag-course
sudo -u postgres psql -d unlearn-rag-course -c "CREATE EXTENSION IF NOT EXISTS vector;"Copy the example environment file and update it with your configuration:
cp .env.example .env.localThen edit .env.local with your settings:
DATABASE_URL=postgresql://postgres:postgres@localhost:5432/unlearn-rag-course
AI_PROVIDER=ollama
AI_MODEL=qwen2.5:14b
EMBEDDING_PROVIDER=voyage
EMBEDDING_MODEL=voyage-4-large
RERANK_PROVIDER=voyage
RERANK_MODEL=rerank-2.5
VOYAGE_API_KEY=your_voyage_api_key_hereFor Unsloth/OpenAI-compatible providers, add:
AI_PROVIDER=unsloth
AI_PROVIDER_BASE_URL=http://localhost:8000
AI_API_KEY=your_api_key_here # Optional, depends on server configUnsloth Model Recommendation: For best results, use Qwen3.6-35B-A3B-GGUF with Q6 quantization. It's overkill but delivers the most accurate and detailed responses.
Note: VOYAGE_API_KEY is required for the reranking stage even if you use EMBEDDING_PROVIDER=ollama for embeddings. The reranker uses Voyage AI's rerank-2.5 model to reorder results from the initial hybrid search.
Available AI Providers:
| Provider | Models | API Key Required |
|---|---|---|
| Ollama | Any local model (e.g., qwen2.5:14b) |
No — runs locally |
| LM Studio | Any local model (e.g., qwen2.5-14b-instruct-mlx) |
No — runs locally |
| Unsloth | Any OpenAI-compatible model (e.g., unsloth/Qwen3-8B-unsloth-bnb-4bit) |
Optional — depends on server config |
| Groq | openai/gpt-oss-120b |
Yes — groq.com |
| DeepSeek | deepseek-v4-flash |
Yes — deepseek.com |
| Anthropic | Any Anthropic model (e.g., claude-sonnet-4-20250514) |
Yes — console.anthropic.com |
| OpenAI | Any OpenAI model (e.g., gpt-4o) |
Yes — platform.openai.com |
Available Embedding Providers:
| Provider | Model | API Key Required |
|---|---|---|
| Voyage AI (default) | voyage-4-large |
Yes — voyageai.com |
| Ollama | Any local embedding model | No — runs locally |
Reranking:
Reranking uses Voyage AI rerank-2.5 regardless of embedding provider. VOYAGE_API_KEY is always required.
Ollama Model Recommendations (≤32GB RAM):
For RAG (answering queries):
| Model | Size | RAM | Quality | Notes |
|---|---|---|---|---|
qwen2.5:32b |
32B | ~20GB | Best | Slowest, best reasoning |
qwen2.5:14b |
14B | ~10GB | Good | Default, good balance |
llama3.1:8b |
8B | ~5GB | Good | Alternative, strong instruction following |
qwen2.5:7b |
7B | ~5GB | Decent | Faster, good for testing |
mistral:7b |
7B | ~5GB | Decent | Fast alternative |
For seeding (generating context during db:seed):
| Model | Size | RAM | Speed | Notes |
|---|---|---|---|---|
qwen2.5:7b |
7B | ~5GB | ~20 min | Minimum for quality context |
qwen2.5:14b |
14B | ~10GB | ~51 min | Default, best context quality |
qwen2.5:32b |
32B | ~20GB | Slowest | Overkill for context generation |
For vector search (embeddings):
| Model | Dimensions | RAM | Notes |
|---|---|---|---|
mxbai-embed-large |
1024 | ~2GB | Recommended, matches pgvector config |
nomic-embed-text |
768 | ~1GB | Lighter, good alternative |
Tip: Use 7b for seeding (~20 min vs ~51 min), then switch to 14b or 32b for queries. Context quality matters — 2b/3b models produce poor context labels that hurt search quality.
# Fast seeding setup (7b is minimum for quality context)
ollama pull qwen2.5:7b
ollama pull mxbai-embed-large
# Configure .env.local for seeding
AI_PROVIDER=ollama
AI_MODEL=qwen2.5:7b
EMBEDDING_PROVIDER=ollama
EMBEDDING_MODEL=mxbai-embed-large
RERANK_PROVIDER=voyage
RERANK_MODEL=rerank-2.5
VOYAGE_API_KEY=your_voyage_api_key_here
# After seeding, switch to quality model for queries
# Edit .env.local:
AI_MODEL=qwen2.5:14bbun db:migrateThis creates all tables: documents, chunks, conversations, messages, message_sources, and rate_limits.
Process the MDN documentation into chunks:
bun chunk-docsThis creates chunks.json from the markdown files in mdn-js-docs/.
bun db:seedThis loads chunks from chunks.json, inserts documents and chunks into the database, and automatically generates embeddings for all chunks. The script uses batch processing with context generation for better search quality. It uses upsert operations — safe to re-run without duplicating data.
If you used bun db:seed, embeddings are already generated. Otherwise, generate embeddings for existing chunks:
bun db:embeddingsThis sends chunks to your configured embedding provider (Voyage AI or Ollama) in batches and stores the resulting vectors in the chunks.embedding column. The script only processes chunks that don't already have embeddings, so it's safe to re-run.
Run automated evaluations of the RAG system:
npm run eval # Run all evaluations
npm run eval:01 # Retrieval accuracy only
npm run eval:02 # Context adherence onlyView detailed results:
npm run eval:view # View all results
npm run eval:view:01 # View retrieval results
npm run eval:view:02 # View context adherence resultsSee evaluation/README.md for setup details.
bun rag-query "What is a closure in JavaScript?"This performs hybrid search and queries the configured AI LLM with retrieved context. Supports --limit flag:
bun rag-query "your question" --limit=10bun devOpen http://localhost:3000 in your browser.
Before using the chat interface, you need to configure your AI provider and API keys:
- Click the Settings icon in the chat header
- Select your AI provider (Groq, DeepSeek, Anthropic, OpenAI, etc.)
- Enter your API key for the selected provider
- Enter your AI model name (e.g.,
openai/gpt-oss-120bfor Groq) - Enter your Voyage API key (required for reranking)
- Enter your embedding model (e.g.,
voyage-4-large) - Click Save
Your settings are stored locally in the browser and used for all chat queries. The CLI scripts (bun rag-query, bun semantic-search) use the .env.local configuration instead.
This project uses Drizzle ORM with PostgreSQL.
The database uses branded types for type-safe IDs. Primary keys use either UUID or text with semantic brand tags to prevent mixing up different ID types at compile time.
| Table | Purpose |
|---|---|
documents |
Source documents (MDN guides) |
chunks |
Document chunks with vector embeddings and BM25 search vectors for hybrid search |
conversations |
Chat conversations |
messages |
Chat messages (user and AI) |
message_sources |
Links between AI messages and source chunks (citations) |
rate_limits |
IP-based rate limiting for the chat API (requests per minute per IP) |
# Generate migrations from schema changes
bun db:generate
# Apply pending migrations
bun db:migrate
# Seed the database with chunks and generate embeddings
bun db:seed
# Generate embeddings for existing chunks
bun db:embeddings
# Debug migration failures
bun db:debug-migrations
# Sync migration journal after manual fixes
bun db:sync-migrations
# Rollback the last migration
bun db:rollbackDatabase scripts live in scripts/db/. Configuration is in drizzle.config.ts.
- AI configuration (
src/config/) — Centralized AI provider and model configuration with support for multiple providers (Ollama, LM Studio, Unsloth, Groq, DeepSeek, Anthropic, OpenAI), embedding providers (Voyage AI, Ollama), and reranking (Voyage AI). User settings from the UI are passed via request headers and override env defaults. - AI providers (
src/lib/aiProviders/) — Provider-specific implementations (e.g., Ollama, LM Studio, and Unsloth via OpenAI-compatible API). - Server logic (
src/lib/server/) — Pure functions for embedding generation, hybrid search (vector + BM25), reranking, context generation, and RAG. Used by both CLI scripts and the Next.js API route. - 3-stage retrieval pipeline (
src/lib/server/search.ts):- Vector + BM25 hybrid search — Combines pgvector similarity search with PostgreSQL full-text search
- Reciprocal Rank Fusion (RRF) — Merges results from both search methods into a single ranked list
- Voyage reranking — Reorders fused results using Voyage AI's
rerank-2.5model for final relevance scoring (gracefully falls back to RRF order on failure)
- Shared constants (
src/lib/shared/) — Configuration like batch sizes and default models, shared between server and client. - API route (
src/app/api/chat/) — Next.js route handler that validates requests and orchestrates the RAG pipeline with tool-based knowledge base access. Reads user AI settings from request headers. - CLI scripts (
scripts/,scripts/db/) — Thin wrappers aroundsrc/lib/server/functions for command-line usage. General scripts inscripts/, database-specific scripts inscripts/db/.
bun dev # Start development server
bun build # Production build
bun type-check # TypeScript type checking
bun lint # Run Biome linter
bun lint:fix # Fix linting issues
bun check-all # Run type-check + lint
bun chunk-docs # Process and chunk documents
bun db:generate # Generate Drizzle migrations
bun db:migrate # Apply database migrations
bun db:seed # Seed database and generate embeddings
bun db:embeddings # Generate embeddings for existing chunks
bun db:debug-migrations # Debug migration failures
bun db:sync-migrations # Sync migration journal
bun db:rollback # Rollback last migration
bun semantic-search "your question" # Search chunks by hybrid search
bun rag-query "your question" # RAG query with LLM response
npm run eval # Run all Promptfoo evaluations
npm run eval:01 # Run retrieval evaluation only
npm run eval:02 # Run context adherence evaluation only
npm run eval:view # View all evaluation results
npm run eval:view:01 # View retrieval results
npm run eval:view:02 # View context adherence resultsFor detailed usage, options, and prerequisites for each script, see scripts/README.md.
- Framework: Next.js 16 (App Router)
- Styling: Tailwind CSS
- Database: PostgreSQL + Drizzle ORM
- Vector Search: pgvector + BM25 full-text search (hybrid search) with Voyage AI reranking
- Embeddings: Voyage AI or local Ollama
- AI/LLM: Vercel AI SDK with Groq, DeepSeek, Anthropic, OpenAI, Ollama, LM Studio, or Unsloth
- Runtime: Bun
- Linting: Biome
This project is based on the RAG Course by VueSchool. It has been heavily modified and extended with multi-provider support, hybrid retrieval, evaluation, and UI improvements.
voyageaiESM build: Thevoyageainpm package has a known ESM import bug (upstream issue). The project usesserverExternalPackagesinnext.config.tsas a workaround.- Dependency advisories: See SECURITY.md for current dependency vulnerability status.
The chat API enforces 20 requests per minute per IP. This is applied at the application level before the request reaches the AI provider. If you exceed this limit, you'll receive a 429 Too Many Requests response. Rate limit state is stored in the rate_limits database table and resets each minute.
This project supports multiple AI providers with different rate limits:
| Provider | Model | RPM | TPM | TPD | Notes |
|---|---|---|---|---|---|
| Groq | openai/gpt-oss-120b |
30 | 12,000 | 100,000 | Free tier — db:seed may hit TPM limit |
| DeepSeek | deepseek-v4-flash |
Check your plan | Check your plan | Check your plan | Higher limits than Groq free tier |
| Voyage AI | voyage-4-large |
Check your plan | Check your plan | Check your plan | Embeddings + reranking |
| Ollama | Any local model | Unlimited | Unlimited | N/A | Runs locally, no rate limits |
| Unsloth | Any model | Depends on server | Depends on server | N/A | OpenAI-compatible API |
Groq free tier limitation: The 12,000 tokens/minute limit can be hit during db:seed or db:generate-contexts since these process many chunks in sequence. If you hit a 429 error, either:
- Wait a minute and retry
- Switch to DeepSeek or Ollama for setup, then use Groq for queries
- Upgrade to a Groq Developer plan for higher limits
If you hit API rate limits, consider switching to a local Ollama model or upgrading to a paid tier.
