Code Nexus
CurriculumHow we teachPricingBlogAboutContact
Start learning
  1. Blog
  2. /
  3. AI engineering

RAG Explained: Retrieval-Augmented Generation

Retrieval-augmented generation explained: why models need your data, how retrieval and grounding work, what usually goes wrong and how to measure it.

By Code Nexus team · 4 min read · Published 2026-09-25 · Updated 2026-09-29 · Reviewed by Code Nexus team

On this page

  • The problem RAG solves
  • Where the idea comes from
  • How it works, step by step
  • The search part is the hard part
  • What usually goes wrong
  • How to measure a RAG system
  • Security and safety
  • RAG and agents
  • Where RAG appears in certifications
  • Learn it by building one

A fluent answer about your company still needs an appropriate source. Retrieval-augmented generation, or RAG, supplies retrieved material at answer time so a model can use it when responding. Retrieval can also find the wrong document, so source checks and evaluation remain part of the work. This guide explains the approach and its limits, within our AI engineering overview.

The problem RAG solves

A model only knows what was in its training data and what you put in its prompt. It cannot know your latest policy, yesterday's incident report or a document written after training finished. And because models are trained to produce fluent text, a gap in knowledge often shows up as a confident wrong answer.

You could paste every relevant document into the prompt, but prompts have limits, and more text is slower, costlier and can distract the model. RAG solves this by finding just the relevant passages first and giving only those to the model.

Where the idea comes from

The name comes from a 2020 paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, by Patrick Lewis and colleagues. They proposed models that combine pre-trained "parametric" memory, meaning what the model learned in training, with "non-parametric" memory, in their case a dense vector index of Wikipedia accessed by a neural retriever. They reported that this produced language that was more specific, diverse and factual than a model that relied on its parameters alone.

Modern systems are built differently, but the principle is the same: look things up, then answer from what you found.

How it works, step by step

Preparing your data (done ahead of time)

  1. Ingest your documents: text, PDFs, and with OCR even scanned images.
  2. Split them into passages, often called chunks.
  3. Index the chunks so they can be searched, commonly by turning each into an embedding, a list of numbers that represents its meaning.

Answering a question (done every time)

  1. Search for the passages most relevant to the question.
  2. Build a prompt containing the question and those passages, with an instruction to answer only from them.
  3. Generate the answer, ideally with references to which passages it used.

The search part is the hard part

The Microsoft AI-103 study guide lists the skills involved: ingesting and indexing documents, images, audio and video; configuring "semantic search, hybrid search, and vector search for grounding"; and enrichment using OCR and layout analysis.

  • Vector search finds passages that mean something similar, even with different words.
  • Keyword search finds exact terms, which matters for names, codes and product identifiers.
  • Hybrid search combines both, and is often the best default.

Most RAG failures are retrieval failures. If the right passage is never found, the model has nothing to work from.

What usually goes wrong

  • Chunks are the wrong size. Too small loses context, and too large dilutes the match.
  • The data is stale or messy. Old versions and drafts get retrieved next to the current policy.
  • The right passage is not found, because the question and the document use different wording.
  • The model ignores the passage or blends it with its own guesses.
  • Nothing is measured, so problems go unnoticed until a customer finds one.

How to measure a RAG system

Build a small evaluation set: real questions with the passage that should answer each. Then measure two things separately:

  1. Retrieval quality: was the right passage in the results?
  2. Answer quality: given the passages, was the answer correct, relevant and grounded in them?

The AI-103 guide names evaluating "fabrications, relevance, quality, and safety" and monitoring "data ingestion quality, search index health, and relevance performance". Separating the two measurements tells you which half to fix.

Security and safety

Retrieved text is untrusted input. A document in your index could contain instructions aimed at the model, which is an indirect form of prompt injection. The OWASP Top 10 for LLM Applications (2025) lists Prompt Injection, Vector and Embedding Weaknesses, and Misinformation among its risks, all of which touch RAG. Practical steps include controlling who can add documents, filtering what is retrieved by the user's permissions, and never letting retrieved text trigger actions without checks. See AI guardrails and safety.

RAG and agents

RAG is often one tool inside a larger system. An agent might decide when to search, which index to search and whether the results are good enough. Understanding RAG first makes agents much easier.

Where RAG appears in certifications

RAG is named directly in the AI-103 outline, and the AWS AIF-C01 exam covers the surrounding decision of whether to prompt, retrieve or fine-tune. The full route is in the certification roadmap.

Learn it by building one

The best way to understand RAG is to build a tiny one over documents you know, then ask questions you know the answers to and see where it fails. That is the approach behind learning by doing. Code Nexus has hands-on episodes on grounding an assistant that invents answers, where you find out why it went wrong and fix it. Create a free account to start with the fundamentals.

Learn AI engineering by doing the work

Code Nexus teaches through interactive story episodes you complete in your browser: real engineering decisions, checked as you go. Create a free account and start with the first episode.

Start Season 1

The first 15 episodes are free. No card needed.

Follow Code Nexus

Frequently asked questions

What does RAG stand for?
Retrieval-augmented generation. The term comes from a 2020 paper by Patrick Lewis and colleagues, which proposed models that combine a pre-trained language model with a searchable store of documents so the model can look up information while generating.
Why not just fine-tune the model on my data?
Fine-tuning changes how a model behaves, and it is a poor way to teach it facts that change. Retrieval lets the model read current documents at answer time and cite them, and updating your data does not require retraining.
Does RAG stop hallucinations?
It reduces them by giving the model source material, but it does not remove them. The model can still ignore the passage or misread it, and retrieval can return the wrong passage. You need to evaluate answers against the sources.
Is RAG on any certification exam?
Yes. Microsoft's AI-103 outline lists implementing retrieval-augmented generation in an application, and configuring semantic, hybrid and vector search for grounding.

Related articles

  • What Is AI Engineering? A Practical Guide

    AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.

  • What Are AI Agents and How Do They Work?

    AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.

  • AI Engineer Skills You Need in 2026

    The skills an AI engineer needs in 2026, from Python and cloud to retrieval, agents, evaluation and safety, with the order to learn them in.

Mentioned in

  • AI Engineer Skills You Need in 2026

    The skills an AI engineer needs in 2026, from Python and cloud to retrieval, agents, evaluation and safety, with the order to learn them in.

  • AI Engineer vs ML Engineer vs Data Scientist

    How AI engineers, machine learning engineers and data scientists differ in daily work, skills and career path, and how to choose which role fits you.

  • AI Guardrails and Safety: What Engineers Must Know

    AI guardrails explained: the main risks to LLM applications, how prompt injection works, and practical controls, with OWASP and NIST as references.

  • What Are AI Agents and How Do They Work?

    AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.

  • What Is AI Engineering? A Practical Guide

    AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.

Sources

  • Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) (accessed 2026-09-25)
  • Study guide for Exam AI-103: Developing AI Apps and Agents on Azure (Microsoft Learn) (accessed 2026-09-25)
  • OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project) (accessed 2026-09-25)

On this page

  • The problem RAG solves
  • Where the idea comes from
  • How it works, step by step
  • The search part is the hard part
  • What usually goes wrong
  • How to measure a RAG system
  • Security and safety
  • RAG and agents
  • Where RAG appears in certifications
  • Learn it by building one
Code Nexus

Practice first. Improvise less later.

We post practical tech tips and the odd 8-BIT opinion. Mostly the tips.

The CPD Group Approved Provider #791172

Learn

  • Curriculum
  • Pricing
  • Create account
  • Sign in

Support

  • Blog
  • Certification guides
  • About Code Nexus
  • How we teach
  • Human help

Legal

  • Terms and Conditions
  • Privacy Policy
  • Cookie Policy
  • Refund and Cancellation
  • Contact: contact@codenexus.co.za

© 2026 Code Nexus. All rights reserved.

Payments secured by Payfast · Billed in ZAR · South Africa