Retrieval-augmented generation
DocQA
Upload a PDF and ask it questions, including about its tables and images. Each answer points back to the page it came from.
Python
LangGraph
Gemini
FastAPI
Streamlit
Techniques: Docling, ChromaDB, BGE reranker, BM25 hybrid, Corrective RAG, RAGAS
View the codeRetrieve
Grade
Rewrite
Answer
- Pick a question to trace it through the graph.
Illustration of the LangGraph control flow, not a recorded run.
The problem
Most PDF chatbots read only the text layer and answer confidently even when the document never says the thing. Real documents hide half their meaning in tables, diagrams and formulas.
How I built it
- 01Docling parses each PDF into typed elements (text, tables, images) with page numbers. Gemini vision describes diagrams and turns formulas into LaTeX, cached so nothing is captioned twice.
- 02Four chunking strategies (fixed, recursive, semantic, structure-aware), with tables and captions always kept whole, embedded into a persistent ChromaDB store.
- 03Retrieval goes wide then narrow: dense top-k, a BGE cross-encoder rerank, optional BM25 hybrid fusion.
- 04A LangGraph corrective-RAG graph grades the retrieved context, rewrites the query once if it isn't enough, and otherwise answers only from the document, falling back to "I don't know based on this document."
Result
Served over FastAPI with a Streamlit chat UI, and evaluated with RAGAS on faithfulness, answer relevancy, context precision and context recall.