AI Guardrails and Safety: What Engineers Must Know
AI guardrails explained: the main risks to LLM applications, how prompt injection works, and practical controls, with OWASP and NIST as references.
By Code Nexus team · 4 min read · Published · Updated · Reviewed by Code Nexus team
A helpful answer in a demo does not establish that an AI system is safe to release. Start by naming the risks, limiting what the system can do and testing the controls. This guide explains common failure modes and practical checks, with OWASP and NIST as references. It connects to the skills in our AI engineering overview.
Why models need guardrails
Models can be tricked, can leak information, can invent facts and can be pushed to produce harmful content. They also behave a little differently each time. Traditional software fails in predictable ways, so tests catch most problems. Model output is variable, so you need extra controls around it, and you need to assume that some input will be hostile.
The reference list: OWASP Top 10 for LLM Applications
The Open Worldwide Application Security Project publishes a widely used list of risks for applications built on large language models. The 2025 version lists:
- LLM01 Prompt Injection
- LLM02 Sensitive Information Disclosure
- LLM03 Supply Chain
- LLM04 Data and Model Poisoning
- LLM05 Improper Output Handling
- LLM06 Excessive Agency
- LLM07 System Prompt Leakage
- LLM08 Vector and Embedding Weaknesses
- LLM09 Misinformation
- LLM10 Unbounded Consumption
You do not need to memorise the numbers. The value is that it turns a vague worry, "is it safe?", into a checklist you can review a design against.
The risks that matter most in practice
Prompt injection
Text in the input tries to change what the model does. Direct injection comes from the user. Indirect injection is hidden in something the model reads: a web page, an email, a document, or even text embedded in an image. Microsoft's AI-103 outline names this last case directly, listing the need to "detect and mitigate indirect prompt injection by using embedded text in images". Assume that anything the model reads could contain instructions from someone else.
Excessive agency
An agent with more access than it needs can cause more harm if it is tricked or simply wrong. See what AI agents are for the design habits that limit this.
Improper output handling
Treating the model's output as safe to run or display. If code executes or renders what the model wrote, an attacker who can influence the model can influence your system.
Sensitive information disclosure and system prompt leakage
The model reveals data it should not, or reveals its own hidden instructions. Do not put secrets in prompts, and filter what data the model can reach.
Misinformation
Fabricated answers presented as fact. Grounding in real sources through RAG helps, but it does not remove the problem, so answers need evaluating.
Unbounded consumption
Loops, retries or very large requests that run up cost or bring the service down. Set limits on tokens, steps and spend.
Practical controls
- Least privilege. Give the system the smallest set of tools and data it needs.
- Approval for risky actions. A human confirms anything that spends money, deletes data or contacts people.
- Input and output filters. Screen for unsafe content, and validate that output matches the format you expect before using it.
- Separate trusted and untrusted content. Mark what came from users and documents so it is not treated as instructions.
- Keep secrets out of prompts and out of anything the model can echo back.
- Limits. Caps on steps, tokens, time and cost.
- Logging and audit trails. Record what the model was given and what it did, so you can investigate.
- Evaluation. Test with adversarial examples, not only friendly ones.
The AI-103 outline covers many of these under responsible AI: safety filters, guardrails, risk detection and content moderation; evaluators and safety evaluations; auditing through trace logging and approval workflows; and governing agent behaviour with oversight modes, constraints and tool-access controls.
The governance side: NIST AI RMF
Safety is also about process. The US National Institute of Standards and Technology released its AI Risk Management Framework on January 26, 2023 to help manage risks to individuals, organisations and society from AI. It is intended for voluntary use and is organised around four functions: Govern, Map, Measure and Manage. Even if you never adopt it formally, it is a useful vocabulary for the questions a team should be asking: who is accountable, what could go wrong, how do we measure it and how do we respond.
Where this appears in certifications
Responsible AI is now a graded topic. Microsoft's AI-103 outline includes it as a major area, and AWS's AIF-C01 gives responsible AI 14% and security, compliance and governance a further 14% of scored content. The certification roadmap shows the route to both.
Learn it by breaking things
Safety is easiest to understand by seeing a system fail. Build a small assistant, feed it a document with a hidden instruction, watch what happens and add the control that stops it. That is how learning by doing applies here, and it is how Code Nexus teaches it: episodes put you in the middle of an incident and ask you to contain it. Create a free account to start with the fundamentals.
Follow Code Nexus
Frequently asked questions
- What is a guardrail in an AI application?
- A control that limits what the system can take in, do or say. Examples include filters on unsafe content, checks on output format, limits on which tools an agent can use, approval steps for risky actions, and monitoring.
- What is prompt injection?
- Text in a model's input that tries to change its behaviour, for example an instruction hidden in a document or web page the model reads. It is the first entry, LLM01, in the OWASP Top 10 for LLM Applications 2025.
- Is there a standard framework for AI risk?
- The US National Institute of Standards and Technology published the AI Risk Management Framework on January 26, 2023. It is voluntary and is organised around four functions: Govern, Map, Measure and Manage.
- Do certifications test AI safety?
- Yes. Microsoft's AI-103 outline includes implementing responsible AI, and the AWS AIF-C01 exam gives responsible AI 14% and security, compliance and governance another 14% of its scored content.
Related articles
- What Is AI Engineering? A Practical Guide
AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.
- What Are AI Agents and How Do They Work?
AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.
- RAG Explained: Retrieval-Augmented Generation
Retrieval-augmented generation explained: why models need your data, how retrieval and grounding work, what usually goes wrong and how to measure it.
Mentioned in
- AI Engineer Skills You Need in 2026
The skills an AI engineer needs in 2026, from Python and cloud to retrieval, agents, evaluation and safety, with the order to learn them in.
- AI Engineer vs ML Engineer vs Data Scientist
How AI engineers, machine learning engineers and data scientists differ in daily work, skills and career path, and how to choose which role fits you.
- RAG Explained: Retrieval-Augmented Generation
Retrieval-augmented generation explained: why models need your data, how retrieval and grounding work, what usually goes wrong and how to measure it.
- What Are AI Agents and How Do They Work?
AI agents explained: how they differ from simple workflows, how tools, memory and retrieval fit in, when to use one, and the risks to design around.
- What Is AI Engineering? A Practical Guide
AI engineering explained: what AI engineers build, how the role differs from data science and ML research, the skills involved, and how to start.
Sources
- OWASP Top 10 for LLM Applications 2025 (OWASP GenAI Security Project) (accessed 2026-09-25)
- AI Risk Management Framework (NIST) (accessed 2026-09-25)
- Study guide for Exam AI-103: Developing AI Apps and Agents on Azure (Microsoft Learn) (accessed 2026-09-25)
- AWS Certified AI Practitioner (AIF-C01) exam guide (AWS Documentation) (accessed 2026-09-25)