Quick summary

  • Wiz observed campaigns targeting LiteLLM, MCP servers, and AI frameworks through RCE, blind prompt injection, and memory credential theft.
  • Model proxies and agents often concentrate provider keys, cloud permissions, and internal connectivity, giving one flaw a large blast radius.
  • Remove unnecessary public AI endpoints, patch gateways, adopt short-lived credentials, and monitor the full input-to-tool-to-egress chain.

What happened

AI infrastructure is being exploited now

Across 90 days of honeypots covering LiteLLM, Flowise, LangChain, Langflow, ChromaDB, Ollama, and related services, Wiz observed sustained attacks with tooling adapted to each product’s internals. Three patterns stood out: exploitation of public MCP servers for remote code execution, blind prompt injection against agent frameworks, and AI-native post-exploitation aimed at credentials and process memory.

The surface is attractive because it concentrates authority. A LiteLLM proxy may hold keys for several model providers, carry cloud IAM permissions, and reach databases or internal APIs through MCP. Agents also accept instructions from external input and act on them, allowing untrusted content to become an instruction channel that operators may not directly see.

Familiar flaws with a new blast radius

The honeypots saw exploitation of an authentication flaw in LiteLLM’s MCP Gateway, where failed token validation could yield an unrestricted empty authorization object, and command injection in MCP server test endpoints. Attackers submitted a fake configuration containing a script that downloaded a cryptominer while returning a valid handshake so the connection test appeared successful.

The first control is to keep control planes, test endpoints, and MCP servers off the public internet unless exposure is required. Gateways must fail closed on authentication errors, commands need strict allowlists, and user input must never be concatenated into a subprocess. Replace durable secrets in environments and memory with short-lived credentials, workload identity, and connector-specific permissions.

Security telemetry should connect input ingestion, tool execution, egress, and resource changes. Vulnerability management must treat AI frameworks like any other production dependency. The honeypot evidence shows attackers already understand AI stacks, so defenders cannot rely on novelty as protection.

Source

Inside 90 days of attacks on AI infrastructure

Why developers should care

Model proxies and agents often concentrate provider keys, cloud permissions, and internal connectivity, giving one flaw a large blast radius.

  1. 1Remove unnecessary public AI endpoints, patch gateways, adopt short-lived credentials, and monitor the full input-to-tool-to-egress chain.