PrismBot RAG Implementation
A production-grade, multi-tenant AI chatbot platform enabling organizations to deploy custom, context-aware assistants across web and WhatsApp channels.
Platform
Web (SaaS) + WhatsApp Business API
Duration
3 Months
The Challenge
Retrieval is where most chatbot deployments break, not generation. A model with a well-tuned prompt still answers wrong when the passages it receives are incomplete, off-topic, or stripped of the context that made them meaningful in the source document. Teams spend weeks tuning prompts and never touch the layer actually producing the errors. PrismBot had to work for organizations deploying assistants over their own material: support handbooks, product manuals, policy documents, internal wikis. In that setting a confidently wrong answer is worse than no answer at all. Naive top-K similarity search kept surfacing fragments severed from their headings, so the model would see a paragraph about renewal windows with no indication of which product it governed. And people phrase questions nothing like the source text: someone asks "how long do I have to send this back" against a document that says "returns must be initiated within 30 days of delivery." Single-vector lookups miss that match on vocabulary alone. Two constraints shaped everything downstream. The platform is multi-tenant, so one organization's embeddings can never surface in another organization's results. A leak here isn't a bug, it's a breach. And it ships to a web widget and a WhatsApp Business number at the same time, which rules out any design where conversation state lives in the browser session.
- Context loss due to naive top-K retrieval
- Fragmented document chunking causing incomplete responses
- No human fallback for failed AI interactions
- Disconnected systems across web and WhatsApp channels
- Risk of cross-tenant data leakage in shared environments
Project Goals
How We Built It
Structured Ingestion Pipeline
Raw documents are noisy. PDFs carry headers, footers, page numbers, and navigation furniture that mean nothing to an embedding model but dilute every vector they land in. We run an LLM preprocessing pass over each document first, stripping that noise and normalizing content into a consistent structure. Chunking then happens against the cleaned structure rather than a fixed character count. Each chunk carries its section and subsection headers inline, so a passage about renewal terms arrives already labeled with the product and policy section it belongs to. The chunk is self-describing even when it's retrieved alone, with none of its neighbors for company. This is the least glamorous stage in the pipeline and the one with the largest effect on output quality. Every improvement further down compounds on top of it.
Advanced Retrieval Strategy
A user asking "how long do I have to send something back" and a document saying "returns must be initiated within 30 days of delivery" share almost no vocabulary. Single-query similarity search misses that match often enough to make an assistant feel unreliable, and the failure is invisible, because the model answers confidently from whatever it did retrieve. We expand each incoming query into multiple rewritten variants, run those searches in parallel against the vector store, and merge the results. Reranking with MMR then drops near-duplicate passages that would otherwise burn context window without adding information, and promotes passages covering different facets of the question. The combination produced roughly four times the answer relevance of single-query top-K retrieval.
Efficient Vector Storage
Vectors live in PostgreSQL with the pgvector extension rather than in a dedicated vector database. Keeping embeddings in the same database as tenant records, documents, and permissions means tenant isolation is enforced by the same constraints protecting everything else, instead of by a second system that has to be kept in sync with the first. An HNSW index handles approximate nearest neighbor search, and every query filters by tenant namespace before the similarity search runs. Retrieval stays under 300ms as the corpus grows. Choosing Postgres over a specialized store costs some raw throughput at the top end. It buys operational simplicity: one database to back up, one place where access rules live, and no window in which a tenant's vectors exist inside a system that doesn't know the tenant boundary.
Multi-Agent Architecture
One prompt doing retrieval, context assembly, and answer generation is quick to write and hard to fix. When answers degrade there's no way to tell which part of the chain caused it, so debugging turns into prompt roulette. We split the work with LangGraph into agents with distinct responsibilities. One owns retrieval and query expansion. One enriches retrieved context and resolves references against conversation history, so "what about the enterprise plan" still works three turns in. One generates the response. One evaluates whether the answer is grounded in the retrieved passages. Each agent is independently testable and independently replaceable. That structure is also what makes escalation possible at all. The grounding check needs somewhere to hand off to when it decides the retrieved context doesn't support an answer, and a monolithic chain has no such seam.
Real-Time Communication & Escalation
Chat runs over WebSockets with Socket.IO, so web and WhatsApp conversations share one server-side state model. A conversation that starts in the website widget doesn't lose its history because the transport changed underneath it. Escalation is a designed path, not an error handler. When the grounding check fails or someone asks for a person, the conversation moves into a human agent queue, and the agent picks it up with the full transcript and the passages the assistant retrieved. They start with context instead of asking the user to repeat themselves. Push notifications and email alerts fire on handoff, so nobody sits in a queue that nobody is watching. Every escalation is also a signal about the corpus. A cluster of handoffs on one topic usually means a document is missing, not that the model is underperforming.
The Outcome
We built PrismBot as a multi-tenant SaaS platform that lets an organization upload its documents, get a trained assistant, and deploy that assistant to a website widget and a WhatsApp Business number from one admin console. Three months from kickoff to production. The build treats retrieval as the primary engineering problem rather than a solved prerequisite. Documents go through an LLM preprocessing pass before anything is embedded, so chunks carry their structural context instead of arriving as orphaned text. Incoming queries are expanded into several phrasings, searched in parallel, and reranked for diversity before a single token reaches the generation model. Vectors live in PostgreSQL with pgvector behind an HNSW index, filtered by tenant namespace at query time. Above that sits a multi-agent layer built on LangGraph. Separate agents own retrieval, context enrichment, response generation, and a grounding check that decides whether the assistant actually has the evidence to answer. When that check fails, or when someone asks for a person, the conversation moves to a human with the full transcript and the retrieved passages attached.
Impact & Results
Key Takeaways
Retrieval quality moves output accuracy further than prompt tuning does
Cleaning and structuring documents before embedding beats tuning the retrieval that reads them
Agents with single responsibilities are debuggable; monolithic prompt chains are not
Human escalation has to be designed into the pipeline, not attached to its failure path
Keeping vectors in the primary database made tenant isolation one problem instead of two
Conclusion & Impact
"Retrieval quality moved output accuracy further than any prompt change we made. Structured ingestion, multi-query retrieval, and MMR reranking each produced gains that held up as the corpus grew, where prompt tuning tended to fix one class of question and quietly break another. The second lesson was architectural. Splitting the pipeline into agents with single responsibilities made the system debuggable, and it made human escalation a designed path rather than something bolted onto the failure case. A retrieval system that can't recognize when it lacks the evidence to answer is a liability no matter how good its retrieval is the rest of the time."
Technology Stack

Ready to achieve similar results?
Let our dedicated engineers bring your vision to life with cutting-edge technology and unparalleled expertise.