Code Nexus
CurriculumHow we teachPricingBlogAboutContact
Start learning
  1. Blog
  2. /
  3. AI engineering

AI Guardrails and Safety: What Engineers Must Know

AI guardrails explained: the main risks to LLM applications, how prompt injection works, and practical controls, with OWASP and NIST as references.

By Code Nexus team · 4 min read · Published 2026-09-25 · Updated 2026-09-29 · Reviewed by Code Nexus team

On this page

  • Why models need guardrails
  • The reference list: OWASP Top 10 for LLM Applications
  • The risks that matter most in practice
  • Practical controls
  • The governance side: NIST AI RMF
  • Where this appears in certifications
  • Learn it by breaking things

A helpful answer in a demo does not establish that an AI system is safe to release. Start by naming the risks, limiting what the system can do and testing the controls. This guide explains common failure modes and practical checks, with OWASP and NIST as references. It connects to the skills in our AI engineering overview.

Why models need guardrails

Models can be tricked, can leak information, can invent facts and can be pushed to produce harmful content. They also behave a little differently each time. Traditional software fails in predictable ways, so tests catch most problems. Model output is variable, so you need extra controls around it, and you need to assume that some input will be hostile.

The reference list: OWASP Top 10 for LLM Applications

The Open Worldwide Application Security Project publishes a widely used list of risks for applications built on large language models. The 2025 version lists:

  1. LLM01 Prompt Injection
  2. LLM02 Sensitive Information Disclosure
  3. LLM03 Supply Chain
  4. LLM04 Data and Model Poisoning
  5. LLM05 Improper Output Handling
  6. LLM06 Excessive Agency
  7. LLM07 System Prompt Leakage
  8. LLM08 Vector and Embedding Weaknesses
  9. LLM09 Misinformation
  10. LLM10 Unbounded Consumption

You do not need to memorise the numbers. The value is that it turns a vague worry, "is it safe?", into a checklist you can review a design against.

The risks that matter most in practice

Prompt injection

Text in the input tries to change what the model does. Direct injection comes from the user. Indirect injection is hidden in something the model reads: a web page, an email, a document, or even text embedded in an image. Microsoft's AI-103 outline names this last case directly, listing the need to "detect and mitigate indirect prompt injection by using embedded text in images". Assume that anything the model reads could contain instructions from someone else.

Excessive agency

An agent with more access than it needs can cause more harm if it is tricked or simply wrong. See what AI agents are for the design habits that limit this.

Improper output handling

Treating the model's output as safe to run or display. If code executes or renders what the model wrote, an attacker who can influence the model can influence your system.

Sensitive information disclosure and system prompt leakage

The model reveals data it should not, or reveals its own hidden instructions. Do not put secrets in prompts, and filter what data the model can reach.

Misinformation

Fabricated answers presented as fact. Grounding in real sources through RAG helps, but it does not remove the problem, so answers need evaluating.

Unbounded consumption

Loops, retries or very large requests that run up cost or bring the service down. Set limits on tokens, steps and spend.

Practical controls

  • Least privilege. Give the system the smallest set of tools and data it needs.
  • Approval for risky actions. A human confirms anything that spends money, deletes data or contacts people.
  • Input and output filters. Screen for unsafe content, and validate that output matches the format you expect before using it.
  • Separate trusted and untrusted content. Mark what came from users and documents so it is not treated as instructions.
  • Keep secrets out of prompts and out of anything the model can echo back.
  • Limits. Caps on steps, tokens, time and cost.
  • Logging and audit trails. Record what the model was given and what it did, so you can investigate.
  • Evaluation. Test with adversarial examples, not only friendly ones.

The AI-103 outline covers many of these under responsible AI: safety filters, guardrails, risk detection and content moderation; evaluators and safety evaluations; auditing through trace logging and approval workflows; and governing agent behaviour with oversight modes, constraints and tool-access controls.

The governance side: NIST AI RMF

Safety is also about process. The US National Institute of Standards and Technology released its AI Risk Management Framework on January 26, 2023 to help manage risks to individuals, organisations and society from AI. It is intended for voluntary use and is organised around four functions: Govern, Map, Measure and Manage. Even if you never adopt it formally, it is a useful vocabulary for the questions a team should be asking: who is accountable, what could go wrong, how do we measure it and how do we respond.

Where this appears in certifications

Responsible AI is now a graded topic. Microsoft's AI-103 outline includes it as a major area, and AWS's AIF-C01 gives responsible AI 14% and security, compliance and governance a further 14% of scored content. The certification roadmap shows the route to both.

Learn it by breaking things

Safety is easiest to understand by seeing a system fail. Build a small assistant, feed it a document with a hidden instruction, watch what happens and add the control that stops it. That is how learning by doing applies here, and it is how Code Nexus teaches it: episodes put you in the middle of an incident and ask you to contain it. Create a free account to start with the fundamentals.

Learn AI engineering by doing the work

Code Nexus teaches through interactive story episodes you complete in your browser: real engineering decisions, checked as you go. Create a free account and start with the first episode.

Start Season 1

The first 15 episodes are free. No card needed.

Follow Code Nexus

Frequently asked questions

What is a guardrail in an AI application?
A control that limits what the system can take in, do or say. Examples include filters on unsafe content, checks on output format, limits on which tools an agent can use, approval steps for risky actions, and monitoring.
What is prompt injection?
Text in a model's input that tries to change its behaviour, for example an instruction hidden in a document or web page the model reads. It is the first entry, LLM01, in the OWASP Top 10 for LLM Applications 2025.
Is there a standard framework for AI risk?
The US National Institute of Standards and Technology published the AI Risk Management Framework on January 26, 2023. It is voluntary and is organised around four functions: Govern, Map, Measure and Manage.
Do certifications test AI safety?
Yes. Microsoft's AI-103 outline includes implementing responsible AI, and the AWS AIF-C01 exam gives responsible AI 14% and security, compliance and governance another 14% of its scored content.

Related articles

  • What Is AI Engineering? A Practical Guide

    AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.

  • What Are AI Agents and How Do They Work?

    AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.

  • RAG Explained: Retrieval-Augmented Generation

    Retrieval-augmented generation explained: why models need your data, how retrieval and grounding work, what usually goes wrong and how to measure it.

Mentioned in

  • AI Engineer Skills You Need in 2026

    The skills an AI engineer needs in 2026, from Python and cloud to retrieval, agents, evaluation and safety, with the order to learn them in.

  • AI Engineer vs ML Engineer vs Data Scientist

    How AI engineers, machine learning engineers and data scientists differ in daily work, skills and career path, and how to choose which role fits you.

  • RAG Explained: Retrieval-Augmented Generation

    Retrieval-augmented generation explained: why models need your data, how retrieval and grounding work, what usually goes wrong and how to measure it.

  • What Are AI Agents and How Do They Work?

    AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.

  • What Is AI Engineering? A Practical Guide

    AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.

Sources

  • OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project) (accessed 2026-09-25)
  • AI Risk Management Framework (NIST) (accessed 2026-09-25)
  • Study guide for Exam AI-103: Developing AI Apps and Agents on Azure (Microsoft Learn) (accessed 2026-09-25)
  • AWS Certified AI Practitioner (AIF-C01) exam guide (AWS Documentation) (accessed 2026-09-25)

On this page

  • Why models need guardrails
  • The reference list: OWASP Top 10 for LLM Applications
  • The risks that matter most in practice
  • Practical controls
  • The governance side: NIST AI RMF
  • Where this appears in certifications
  • Learn it by breaking things
Code Nexus

Practice first. Improvise less later.

We post practical tech tips and the odd 8-BIT opinion. Mostly the tips.

The CPD Group Approved Provider #791172

Learn

  • Curriculum
  • Pricing
  • Create account
  • Sign in

Support

  • Blog
  • Certification guides
  • About Code Nexus
  • How we teach
  • Human help

Legal

  • Terms and Conditions
  • Privacy Policy
  • Cookie Policy
  • Refund and Cancellation
  • Contact: contact@codenexus.co.za

© 2026 Code Nexus. All rights reserved.

Payments secured by Payfast · Billed in ZAR · South Africa