AiVibe

AI & Machine Learning

Retrieval-Augmented Generation (RAG)

Retrieval-augmented generation (RAG) improves a language model's answers by first retrieving relevant passages from a trusted knowledge source, such as company documents, and supplying them in the prompt so that the model answers from that information.

Documents are split into chunks, converted into embeddings and stored in a vector index, often alongside a keyword index. When a user asks a question, the system retrieves the most relevant chunks, optionally re-ranks them, and inserts them into the prompt with instructions to answer from the supplied context and cite sources. The approach was described in a 2020 paper by Lewis and colleagues at Facebook AI Research.

RAG lets assistants answer from current, organisation-specific information, such as maintenance manuals, standard operating procedures, quality records and support tickets, without retraining the model. Updating the knowledge base updates the answers, and citations let users verify the source.

Answer quality depends heavily on retrieval: poor chunking, missing documents or weak search lead to wrong or incomplete answers, and the model may still add claims beyond the context. Evaluation measures retrieval recall, faithfulness to sources and answer correctness. Access control must ensure users only retrieve documents they are entitled to see, and retrieved content can carry indirect prompt injection.

Key points

Where AiVibe comes in

AiVibe delivers chatbots and virtual assistants with RAG, MCP tools and voice.

Explore AiVibe’s work in AI & Machine Learning →

Related terms

Terms that refer to Retrieval-Augmented Generation (RAG)

Ask AiMuruga can explain Retrieval-Augmented Generation (RAG) for your plant, product or security programme, and draw how it fits.