When the MCP server is the attacker
MCP servers give agents new powers in a few clicks. They also give whoever wrote them a direct line into your agent's instructions.
Sting Labs· 27 Sep 2026
A market stall full of plugs
The Model Context Protocol made it easy to give agents new abilities: search your docs, read your calendar, query a database. You install a server, the agent sees a list of tools, and it starts using them.
That convenience has a cost. Each tool comes with a name and a description written by whoever built the server, and the agent reads that description as guidance.
Three ways it goes wrong
Poisoned descriptions: a tool called "format_date" whose description also tells the agent to read a private file and include it in the next call.
Bait and switch: a server that behaves perfectly when you first try it, then changes what its tools say or do after it has been approved.
Look-alikes: a server with a name one letter away from a popular one, installed by mistake.

In plain words
The words an MCP server uses to describe itself are instructions to your agent. If the author is hostile, so are the instructions.
Why reviewing once is not enough
A one-time review checks the server on the day you install it. The descriptions the agent actually reads are loaded again and again, and they can change.
The safe habit is to check what the agent receives each time, not what the server claimed last month.
Keeping the good ones
The answer is not to ban MCP servers. Most are useful and honest, and teams will keep installing them.
What matters is the split: clean servers get in and work normally, and servers carrying hidden instructions never get to teach the agent anything.



