MalwareTips News Malware authors are planting instructions to mislead AI security tools

Would you trust an AI-generated “safe” verdict if your antivirus still warned about the file?

  • No, I would trust the warning

    Votes: 0 0.0%
  • I would seek a second opinion

    Votes: 1 100.0%
  • Yes, I might trust the AI

    Votes: 0 0.0%
  • I am not sure

    Votes: 0 0.0%
  • I do not use AI security tools

    Votes: 0 0.0%

  • Total voters
    1

News Now

Happening Now
Verified
MalwareTips-news-192.jpg

Image: Cisco Talos

Malware developers are embedding plain-text instructions that try to persuade AI analysis tools to ignore or misclassify malicious files. The immediate concern is mainly for security products and analysts using language models; conventional malware detection remains unaffected by this technique.


What researchers found​

Cisco Talos and CAIRN classify this tactic as “AI-analysis evasion.” It targets systems that extract text from a suspicious file and send it to a large language model for classification, triage or reverse-engineering help.

Researchers traced the approach across 84 samples from four confirmed malware families—FRUITSHELL, PLOTSAFE, HOLLOWCLAD and MANTLEMAZE—collected between January 2025 and July 2026.

  • FRUITSHELL included comments telling AI that the script was harmless prime-number software, even though no such code existed.
  • HOLLOWCLAD placed refusal instructions in seven AI chat-template formats, hoping one would match the analyzer’s format.
  • MANTLEMAZE used fabricated government contracts, certifications and ownership claims to try to trigger model guardrails.

Simple tricks sometimes worked better​

Talos tested extracted instructions with five locally run language models. Each malicious file was compared with and without the planted text, and the runs measured whether the model shifted toward a benign, suspicious or malicious verdict.

Direct instructions to ignore a sample were the most consistently effective approach, while more elaborate attempts often had little impact or made models more suspicious. The best techniques steered results in the attacker’s favor in about 35% of test runs.

How defenders can reduce the risk​

The planted instructions must remain readable as plain text, giving security teams a stable way to detect them. Commands aimed at an analyzer—such as demands to stop analysis or accept copyright claims—should themselves be treated as suspicious evidence.

  • Build AI analysis tools so text extracted from a file is always treated as untrusted evidence, never as an instruction.
  • Keep conventional antivirus and endpoint detection enabled because this tactic targets an AI-assisted analysis layer rather than those core detection methods.
  • Investigate conflicting results instead of trusting an AI verdict that disagrees with other security alerts.
 

Recently browsing

Members who viewed this thread in the last 5 minutes

Back
Top