Bengali Book-Based RAG Chatbot Development by Rabiul IslamBengali Book-Based RAG Chatbot Development by Rabiul Islam

Bengali Book-Based RAG Chatbot Development

Rabiul Islam

Rabiul Islam

📚 āĻŦāĻŋāĻļā§āĻŦ⧇āϰ āωāĻĒāĻžāĻĻāĻžāύ — Knowledge Base Chatbot

A Bengali book-based Retrieval-Augmented Generation (RAG) chatbot built using LangChain, multilingual embeddings, FAISS, Groq, and Streamlit.
The chatbot answers questions using only the selected Bengali book “āĻŦāĻŋāĻļā§āĻŦ⧇āϰ āωāĻĒāĻžāĻĻāĻžāĻ¨â€ and provides the relevant chapter and source URL with each answer.

đŸŽ¯ Project Objective

The goal of this project is to build a Knowledge Base Chatbot with a Vector Database that can:
Crawl a complete Bengali book from Wikisource
Clean and preprocess the collected text
Split the book into meaningful chunks
Generate multilingual embeddings
Store embeddings in a FAISS vector database
Retrieve relevant book passages
Generate answers using an LLM
Provide chapter and source information
Avoid using outside knowledge for book-specific questions
Clearly respond when information is not available in the selected book

📖 Knowledge Source

Selected Book

Book: āĻŦāĻŋāĻļā§āĻŦ⧇āϰ āωāĻĒāĻžāĻĻāĻžāύ Author: āĻļā§āϰ⧀āϚāĻžāϰ⧁āϚāĻ¨ā§āĻĻā§āϰ āĻ­āĻŸā§āϟāĻžāϚāĻžāĻ°ā§āϝ Publication Year: 1952 Publisher: Visva-Bharati Source: Bengali Wikisource
Source URL:
The selected book is a completed Bengali prose book containing six main chapters.

Chapters

āĻ…āϪ⧁, āĻĒāϰāĻŽāĻžāϪ⧁
āχāϞ⧇āĻ•ā§â€ŒāĻŸā§āϰāύ āĻ“ āĻĒā§āϰ⧋āϟāύ
āĻĒāϜāĻŋāĻŸā§āϰāύ, āύāĻŋāωāĻŸā§āϰāύ, āύāĻŋāωāĻŸā§āϰāĻŋāύ⧋, āĻŽāĻŋāϏ⧋āĻŸā§āϰāύ
āωāĻĒāĻžāĻĻāĻžāύ⧇āϰ āĻĒā§āϰāĻ•ā§ƒāϤāĻŋ
āĻļāĻ•ā§āϤāĻŋ āĻ“ āϤāĻĄāĻŧāĻŋā§Ž
āωāĻĒāϏāĻ‚āĻšāĻžāϰ

🧠 RAG Architecture


🔍 Retrieval Configuration

The final chunking configuration is:

The processed book produces:

The chatbot uses:

The retriever uses FAISS similarity search with:

This configuration was selected after testing different chunk sizes.

📊 Retrieval Evaluation

The project includes 10 test questions:
9 answerable questions
1 no-answer question
For the 9 answerable questions, the expected chapter was found within the top-10 retrieved results.

Result


Note: This 100% result represents retrieval coverage at top-10, not 100% end-to-end answer accuracy.

The no-answer test is separately used to evaluate whether the chatbot can avoid answering questions outside the selected book.

đŸ—‚ī¸ Project Structure


venv/ and .env are intentionally excluded from Git because they are listed in .gitignore.

âš™ī¸ Installation

1. Clone the repository


2. Create a virtual environment

Windows:

Activate the environment:

3. Install dependencies


🔐 Environment Variables

Create a .env file in the project root:

Never commit the real API key to GitHub.
The project .gitignore contains:

so the API key remains local.

🚀 Run the Chatbot

Activate the virtual environment:

Run the Streamlit application:

Open the application in your browser:

đŸ§Ē Test Questions

The project contains 10 test questions.

1. āĻĒāϰāĻŽāĻžāϪ⧁ āϕ⧀?

2. āχāϞ⧇āĻ•ā§āĻŸā§āϰāύ āϕ⧀?

3. āĻĒā§āϰ⧋āϟāύ āϕ⧀?

4. āύāĻŋāωāĻŸā§āϰāύ āϕ⧀?

5. āĻĒāϜāĻŋāĻŸā§āϰāύ āϕ⧀?

6. āĻŽāĻŋāϏ⧋āĻŸā§āϰāύ āĻŦāĻž āĻŽā§‡āϏāύ āϏāĻŽā§āĻĒāĻ°ā§āϕ⧇ āĻŦāχāϟāĻŋāϤ⧇ āϕ⧀ āĻŦāϞāĻž āĻšā§Ÿā§‡āϛ⧇?

7. āĻŽā§ŒāϞāĻŋāĻ• āĻĒāĻĻāĻžāĻ°ā§āĻĨ⧇āϰ āĻĒāϰāĻŽāĻžāϪ⧁āϗ⧁āϞ⧋āϰ āĻŽāĻ§ā§āϝ⧇ āϕ⧀ āĻĒāĻžāĻ°ā§āĻĨāĻ•ā§āϝ āφāϛ⧇?

8. āĻļāĻ•ā§āϤāĻŋ āĻ“ āϤ⧜āĻŋā§Ž āϏāĻŽā§āĻĒāĻ°ā§āϕ⧇ āĻŦāχāϟāĻŋāϤ⧇ āϕ⧀ āφāϞ⧋āϚāύāĻž āĻ•āϰāĻž āĻšā§Ÿā§‡āϛ⧇?

9. āĻĒā§āϰāĻžāωāĻŸā§‡āϰ āĻŽāϤ āϕ⧀ āĻ›āĻŋāϞ?

10. āĻŦāĻžāĻ‚āϞāĻžāĻĻ⧇āĻļ⧇āϰ āĻŦāĻ°ā§āϤāĻŽāĻžāύ āϜāύāϏāĻ‚āĻ–ā§āϝāĻž āĻ•āϤ?

Question 10 is intentionally outside the selected book and is used as a no-answer test.
Expected behavior:

đŸ›Ąī¸ Hallucination Control

The chatbot is instructed to answer book-specific questions only from the retrieved book context.
The prompt enforces the following rules:
Use only the retrieved book context.
Do not use outside knowledge.
Use relevant information even if the wording differs from the question.
Answer in Bengali.
Include the relevant chapter and source.
If the information is not available in the selected book, clearly say:

This prevents the chatbot from behaving like a general-purpose knowledge engine.

🌐 Multilingual Embeddings

The source material is Bengali, so an English-only embedding model would not be appropriate for semantic retrieval.
The project uses:

This model supports multilingual semantic representations and works with the Bengali book content.

đŸ—„ī¸ Why FAISS?

FAISS is used as the vector database because it provides efficient similarity search over embedding vectors and is convenient for a local RAG application.
The vector store is generated from the processed book chunks.

đŸˇī¸ Metadata Preservation

Each document chunk preserves important metadata such as:

This allows the chatbot to identify where retrieved information came from and provide source references with answers.

🧩 Main Components

crawler.py

Downloads the selected book chapters from Bengali Wikisource.

processor.py

Cleans raw scraped text and preserves important metadata.

chunker.py

Splits cleaned documents into overlapping text chunks.
Final configuration:

embedding.py

Generates multilingual embeddings and creates the FAISS vector database.

retriever.py

Loads the FAISS database and retrieves the most relevant book chunks.

qa.py

Combines retrieval with the Groq LLM and generates Bengali answers based only on retrieved context.

app.py

Provides the Streamlit chatbot interface.

🧰 Technologies Used

Technology Purpose Python Core programming language LangChain RAG pipeline LangChain Community FAISS and supporting integrations LangChain HuggingFace Embedding integration LangChain Text Splitters Document chunking BeautifulSoup Web scraping Requests HTTP requests Sentence Transformers Multilingual embeddings FAISS Vector similarity search Groq LLM inference Streamlit Web interface python-dotenv Environment variable management

📌 Important Design Decisions

Multilingual embedding model

The source material is Bengali, so multilingual embeddings were selected instead of an English-only embedding model.

Chunk size

Different chunking configurations were tested.
The final configuration:

produced 103 chunks and achieved 9/9 expected-chapter retrieval coverage in the current test set.

Top-k retrieval

The retriever uses:

This was chosen because some questions, such as āĻĒā§āϰ⧋āϟāύ āϕ⧀?, did not consistently retrieve the expected chapter within a smaller top-k value but were found when using top-10 retrieval.

🔮 Future Improvements

Possible improvements include:
Bengali-specific embedding models
Hybrid keyword + semantic retrieval
Reranking retrieved documents
Better no-answer detection
Streaming LLM responses
Improved source citation UI
Chunking strategy comparison
Embedding model comparison
Automated retrieval evaluation
End-to-end answer evaluation
Deployment using Streamlit Cloud or another hosting platform

👨‍đŸ’ģ Author

Rabiul Islam
AI/ML Developer | Computer Vision | Generative AI | AI Agents
GitHub:
Project Repository:

📜 Disclaimer

This project is an educational RAG application built around the selected Bengali book “āĻŦāĻŋāĻļā§āĻŦ⧇āϰ āωāĻĒāĻžāĻĻāĻžāĻ¨â€.
The chatbot is designed to answer questions using the selected knowledge source and should not be treated as a general-purpose factual search engine.

⭐ Bonus — Chunking Strategy Comparison

Two different chunking strategies were evaluated to measure their effect on retrieval performance.

Approaches Tested

Strategy Chunk Size Chunk Overlap Total Chunks Strategy A 700 150 135 Strategy B 1000 200 103
Both strategies used the same multilingual embedding model:
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
The same 9 answerable test questions and the same top-10 retrieval setting were used for both strategies.

Evaluation Method

For each test question, the FAISS vector store retrieved the top 10 chunks.
A retrieval was counted as a hit when the expected chapter appeared in at least one of the top 10 retrieved chunks.
Hit Rate = Correct Retrievals / Total Answerable Questions × 100

Results

Strategy Correct Retrievals Hit Rate 700 / 150 7 / 9 77.8% 1000 / 200 9 / 9 100.0%

Result

The 1000 / 200 configuration achieved a 100.0% retrieval hit rate on the 9 answerable test questions, compared with 77.8% for the 700 / 150 configuration.
Therefore, the 1000 / 200 configuration was selected for the final RAG pipeline based on this evaluation.

Note: This hit-rate evaluation measures retrieval coverage only. It does not represent the accuracy of the final LLM-generated answers and should not be interpreted as a universal result for other datasets or books.

Like this project

Posted Sep 18, 2026

Developed a RAG chatbot for a Bengali book using LangChain and FAISS.