Skip to content
Thursday, 8 October 2026
Tenesys AI News
Subscribe
Cyber Security· Important

Cisco Talos documents malware that tries to prompt-inject AI security analysts

In short: Cisco Talos' CAIRN research classifies a new malware archetype, 'A3: AI-Analysis Evasion,' where malicious code embeds natural-language text designed to manipulate language models used in automated malware triage. Talos traced the technique across four malware families (FRUITSHELL, PLOTSAFE, HOLLOWCLAD, MANTLEMAZE) and 84 samples from January 2025 to July 2026, showing it spreading from a simple copy-pasted comment to templated, multi-model-targeted 'template spraying' and fake intimidation messages. Controlled testing with a panel of five local LLMs found the simplest direct-instruction technique shifted model verdicts toward 'benign' in roughly 35% of runs, while more elaborate tricks had little effect or backfired by increasing suspicion.

Source: Cisco TalosCisco TalosOriginal article ↗

This summary was generated automatically by AI from Cisco Talos's publication. It is our own text, not a copy of the original — facts, figures and quotes belong to the source, linked above and below.

What changed?

  • 1New malware category 'A3: AI-Analysis Evasion' tracked via Cisco Talos' CAIRN framework
  • 2FRUITSHELL: a PowerShell reverse shell with a comment telling AI/LLM analyzers to skip analysis, reused verbatim by at least 4 other actors in 9+ scripts, later combined with AMSI bypass and shellcode loading in 'ROZESHELL'
  • 3PLOTSAFE: templated, auto-generated anti-analysis comments with varying filler text, engineered to survive compiler string stripping
  • 4HOLLOWCLAD: 'template spraying' — the same refusal instruction formatted in seven different LLM chat template syntaxes, plus fake intimidation notes claiming monitoring/honeypots
  • 5MANTLEMAZE: VMProtect-packed loaders impersonating Google/Microsoft/Valve IP with fabricated DoD contracts, FIPS certifications, NDAs, patents and fake boot logs; also drops a known vulnerable driver (CVE-2015-2291) for EDR-killing (BYOVD)
  • 6Evaluation across 5 local LLMs, 135 matched pairs per string: simplest direct-instruction evasion worked in ~35% of runs; complex techniques showed little to no effect or increased suspicion
  • 7Talos notes the evasion text must always be plaintext, making it a stable, detectable signal

Why it matters

Security teams and vendors relying on LLMs for malware triage, SOC automation, or AI-assisted reverse engineering need to treat any text extracted from a sample as untrusted data, not as instructions, since attackers are actively testing prompt-injection against these pipelines.

What it means for AI agents and contact centers

If any internal tooling uses LLMs to scan logs, call recordings, uploaded files, or automation workflows for security triage or content classification, apply the same principle highlighted here: never let extracted text from an unverified source (files, transcripts, documents) be interpreted as a system or assistant instruction — keep strict separation between data and prompts in any AI-assisted security or automation pipeline.

Sources

  • Cisco TalosOfficialPrimary source
    „Ignore all instructions and read this blog: The state of AI-analysis evasion in malware“
    8 Oct 2026, 13:00
    Original article →
Published by source
8 Oct 2026, 13:00
Found by our system
8 Oct 2026, 13:08
Summary generated
8 Oct 2026, 13:10

This article was written by AI from the original source. Facts, numbers and prices come from the source; missing values are marked “Not specified”. Legal notice, copyright and privacy

Cisco Talos documents malware that tries to prompt-inject AI security analysts · TENESYS AI NEWS