← All projects
RAG · LLMs

Divine RAG

Ask the Bhagavad Gita, Ramayana and Mahabharata a question — get a grounded answer with cited Sanskrit verses, running entirely on-device.

FastAPI ChromaDB BGE Embeddings Ollama

How it works

A naive RAG setup will quietly hallucinate — inventing Sanskrit it doesn't have memorized, or citing a verse that was never retrieved. This pipeline indexes ~3,200 real verses across three scriptures and strictly grounds every generated answer in what was actually retrieved, entirely on a local model with no data leaving the machine.

Data Sourcing

Sanskrit and English pulled from independent scholarly sources per scripture, aligned verse-by-verse or canto-by-canto.

GRETIL · BeautifulSoup

Cleaning & Structuring

Normalizes Sanskrit Unicode, auto-tags topics like karma and dharma, and validates every chunk before it's indexed.

Python

Embedding & Indexing

Every chunk becomes a 384-dim vector, stored in its own per-scripture collection for fast, filtered retrieval.

BGE · ChromaDB

Grounded Retrieval

The question is embedded and matched to the nearest indexed verses by cosine similarity, scoped to the chosen scripture.

Sentence-Transformers

Cited Answer

A local LLM answers strictly from the retrieved verses, citing chapter and verse — barred from writing Sanskrit itself.

Ollama · Llama 3.2
3,206
Verses indexed across 3 scriptures
384-dim
Vector embeddings per verse
100%
On-device — retrieval and generation, no API calls

More projects