CydraLabs

Resources · Threat Lab

How AI agents are attacked, and what controls them

Briefings on the techniques we design CydraLabs controls against. Each lists the relevant controls and their current status. Examples are illustrative, not incident reports.

Techniques

Threat briefings

  • MCP tool poisoning

    A tool's description, schema or output carries instructions that steer the agent: calling another tool, widening a query or sending data somewhere it should not go.

    Illustrative example: A newly added MCP server describes its search tool with hidden text telling the agent to forward results to an external address.

    Relevant controls

    • MCP server and tool discovery, so new tools are visibleBeta
    • Deny-by-default tool allowlists and pre-execution policyAvailable
    • Prompt-injection monitoring on tool parametersBeta
  • Identity abuse

    Agents run with shared, long-lived or over-privileged credentials, or with delegated user authority nobody reviewed. A stolen token or confused agent then acts with far more reach than intended.

    Illustrative example: An agent's static API key is copied from a repository and replayed from outside the organisation.

    Relevant controls

    • Signed, five-minute workload tokens with replay protectionAvailable
    • Delegation mapping in the Security GraphAvailable
    • Per-execution scoped credential handlesBeta
  • Memory attacks

    Persistent agent memory is poisoned with instructions or false facts that influence later sessions, possibly for other users.

    Illustrative example: A support conversation plants a note in long-term memory that tells the agent to skip verification for a particular account.

    Relevant controls

    • Persistent memory as an Agent Risk Score factorAvailable
    • Policy and approval on the actions memory would influenceAvailable
    • Memory inspection controlsRoadmap
  • Excessive agency

    An agent can take consequential actions — payments, deletions, code merges, access changes — without a human in the loop or a limit on scope.

    Illustrative example: A finance agent is asked to pay an invoice and is technically able to pay any amount to any supplier.

    Relevant controls

    • Allow, deny or require-approval decisions before executionAvailable
    • Single-use approval tokens bound to the exact requestAvailable
    • Agent kill switchAvailable

Incident analyses

Coming soon

Analyses of real agent incidents, using public information and our own testing, will be published here. None are published yet. If you have found a vulnerability in a CydraLabs service, please report it.

Report a security issue