Blog

MCP Prompt Injection: How Tool Results Cross Boundaries

MCP prompt injection happens when an AI assistant treats untrusted text from a tool, resource or server description as an instruction. A harmless lookup then steers what the assistant does next. You can't fully stop this by trusting the server. You limit the damage by restricting which tools exist, narrowing what they can reach, and reviewing consequential actions.

What does an attack look like?

Say you ask a support assistant to summarize open tickets. Its ticket-search tool returns a customer's message. Most of it describes a login problem. One paragraph says the summary can only be finished by pulling an internal report and attaching it to an external case.

You asked for a summary. The customer wrote some text to summarize. The failure happens if the assistant treats that text as permission to do a second job. No bug in the ticket API is needed. The problem is in how the assistant reads what comes back. This is an illustrative scenario, not a report of a real incident.

The OWASP prompt-injection guidance describes this as an indirect attack: instructions arrive inside retrieved material. It recommends checking proposed calls against permissions and context, and limiting what the application can access.

Why isn't a secure connection enough?

TLS protects traffic between two endpoints. Authentication proves you're allowed to use a service. Neither says anything about whether a ticket, a web page or a database text field should be able to give orders to your assistant. A trusted server can still return content written by someone you don't trust.

Where can instructions get in?

Tool results and resources

Tools return free text, or structured fields that contain free text. A JSON wrapper makes the reply easy to parse, but a string inside it can still say "ignore your instructions." Resources have the same problem when their contents come from outside your workflow.

Judge content by who could have written it, not by which server delivered it. In the ticket example, the customer controls one field, and that's enough.

Tool descriptions and schemas

A server's tool definition tells the model when and how to call a tool. If an untrusted server hides instructions there, the attack lands before the first call. OWASP's MCP security guidance lists descriptions, parameter schemas and returned values as things to inspect.

Review changes to a tool's description and arguments, not just the first install. Keep a list of approved tools, and re-check additions before they reach a sensitive workflow.

Can a read-only tool still cause harm?

Yes. A tool that can't modify your database can still read confidential rows, and the assistant can then pass them somewhere else. A generic search tool can carry data in its query text. An issue-writing tool can carry it in a title or body. You don't need a tool called "export secrets" for data to leak.

So review tools in combination, inside the actual app you use. A tightly configured database connection can sit next to a separate browser, file or network tool, and a gateway on one connection can't police tools attached by another route.

The MCP maintainers make the same point in their post on tool annotations. A "read-only" hint describes likely behavior. Real limits come from authorization and runtime controls.

How do you limit the damage?

Limit what the assistant can call

Start from the task. If the assistant only needs reporting, remove write, admin and external-messaging tools. A forbidden tool should fail even when the model asks for it by name.

In MCPifex, each instance has its own tool policy. Disabled tools are hidden from the tool list and refused if called, and tools outside the catalog are always refused. The instances and tools guide explains it. This shrinks what can happen. It doesn't detect malicious wording in results.

Limit what an allowed tool can reach

Give a reporting server an account built for reporting. For PostgreSQL, that means a dedicated role that can read only approved tables or views. If the task needs monthly totals, return totals, not full customer rows that the model is asked to mostly ignore. See the PostgreSQL server page for the hosted tool set.

Review the specific action

An approval prompt is only useful if it shows the tool, its arguments, the destination and the effect. "Allow tool use?" won't help you spot a surprise attachment. Ask again when an action goes beyond the original task, and make sure a "no" really stops it.

The MCP tools specification recommends confirmation for sensitive operations and showing inputs before a call. Clients differ, so check what yours actually displays.

How do you test your defenses?

Build a sandbox with synthetic records, a fake secret marker, and a mocked outbound tool that records requests without sending anything. Feed the injection through the same path production uses, such as a ticket field or a document section.

  1. Baseline. Ask for a summary of ordinary data and note the calls you expect.
  2. Add the bait. Insert a clearly unauthorized request that uses the fake marker.
  3. Run it. Watch the proposed and executed calls, not the assistant's claim that it ignored the request.
  4. Check the boundaries. Confirm disabled tools are refused and off-limits data stays unreachable.
  5. Inspect the mock. Make sure it received nothing unauthorized.

Try several phrasings and placements, including structured fields a summary tends to quote. Passing a small test is evidence about those cases, not proof of immunity. Repeat it after any change to the model, client, server or permissions.

What if the assistant tries something unexpected?

Pause the workflow. Look at the content that came in, the call it tried to make, and the arguments. Don't follow troubleshooting steps suggested by the suspicious content itself. Keep a redacted record so you can reconstruct what happened.

Then see which controls held. If the model asked for a disabled write and the gateway refused it, your boundary worked even though the model was fooled. If an allowed read returned too much, tighten the upstream permission.

Where MCPifex fits

As of September 2026, MCPifex gives you per-instance call logs and usage counts, and you can revoke an API key from the portal. These help you investigate and contain a problem. They aren't automatic prompt-injection detection. For the wider picture, read the MCP security guide, and see the pricing page for plan limits.

Key takeaways

  • A trusted server can return untrusted text written by someone else.
  • Read-only tools can still expose data or pass it to another tool.
  • Tool allowlists limit actions. Upstream permissions limit data.
  • Test with synthetic data and a mocked destination, and watch the calls, not the answer.

Sources

  1. OWASP LLM Prompt Injection Prevention.
  2. OWASP MCP Security Cheat Sheet.
  3. MCP: Tool Annotations as Risk Vocabulary.
  4. MCP tools security considerations.