Prompt Injection
Prompt injection is an attack in which crafted text causes a language model to ignore its intended instructions and follow an attacker's instead, either typed directly by a user or hidden in content the model processes, such as web pages, emails or documents.
Language models process instructions and data in the same stream of text and cannot reliably tell them apart. In direct injection, a user writes instructions designed to override the system prompt. In indirect injection, malicious instructions are embedded in retrieved documents, websites, emails, tool outputs or images, and the model acts on them when it reads that content. Jailbreaking, which tries to bypass a model's safety training, is a closely related technique.
The risk grows with agents and assistants that have tools: an injected instruction might cause data exfiltration, unauthorised actions, misleading answers or manipulation of downstream systems. Prompt injection is ranked first in the OWASP Top 10 for LLM Applications.
There is no complete fix. Defences include least-privilege tool access, human confirmation of sensitive actions, separating and labelling untrusted content, input and output filtering, restricting where data can be sent, and monitoring. Systems should be designed so that a successful injection has limited impact, and red teaming should test for it regularly.
Key points
- Crafted text overrides a model's intended instructions
- Indirect injection hides instructions in documents, web pages or tool output
- Ranked first in the OWASP Top 10 for LLM Applications
- No complete fix exists, so impact must be limited by design
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.
Related terms
- AI GuardrailsAI & Machine Learning
- AI AgentAI & Machine Learning
- Retrieval-Augmented Generation (RAG)AI & Machine Learning
- Function Calling (Tool Use)AI & Machine Learning
- OWASP Top 10Cybersecurity & Compliance
- Model Context Protocol (MCP)AI & Machine Learning