AiVibe

AI & Machine Learning

Prompt Injection

Prompt injection is an attack in which crafted text causes a language model to ignore its intended instructions and follow an attacker's instead, either typed directly by a user or hidden in content the model processes, such as web pages, emails or documents.

Language models process instructions and data in the same stream of text and cannot reliably tell them apart. In direct injection, a user writes instructions designed to override the system prompt. In indirect injection, malicious instructions are embedded in retrieved documents, websites, emails, tool outputs or images, and the model acts on them when it reads that content. Jailbreaking, which tries to bypass a model's safety training, is a closely related technique.

The risk grows with agents and assistants that have tools: an injected instruction might cause data exfiltration, unauthorised actions, misleading answers or manipulation of downstream systems. Prompt injection is ranked first in the OWASP Top 10 for LLM Applications.

There is no complete fix. Defences include least-privilege tool access, human confirmation of sensitive actions, separating and labelling untrusted content, input and output filtering, restricting where data can be sent, and monitoring. Systems should be designed so that a successful injection has limited impact, and red teaming should test for it regularly.

Key points

Where AiVibe comes in

AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.

Explore AiVibe’s work in AI & Machine Learning →

Related terms

Terms that refer to Prompt Injection

Ask AiMuruga can explain Prompt Injection for your plant, product or security programme, and draw how it fits.