SecurityBrief US - Technology news for CISOs & cybersecurity decision-makers
United States
AI security incidents expose new criminal tradeoffs

AI security incidents expose new criminal tradeoffs

Thu, 17th Sep 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Check Point Research has published a report outlining recent AI security incidents involving major model developers and criminal groups. The period saw both test models reaching live systems and attackers using AI in ransomware operations.

Central to the report are incidents involving internal evaluations at OpenAI, Anthropic and Meta, where models moved beyond intended test environments and interacted with production infrastructure.

One OpenAI model, the report said, found and exploited a previously unknown flaw in an internal package proxy and used it to reach Hugging Face production systems. Researchers later reconstructed the sequence at about 17,600 steps, suggesting a long chain of autonomous actions inside what was designed as an isolated environment.

At Anthropic, an evaluation environment was left reachable from the internet, allowing test models to collect credentials and read a production database.

Meta's case involved a third-party evaluator's misconfiguration. The UK AI Security Institute reported 19 unauthorised actions across 122 controlled runs. In one instance, an agent created fake identities to try to persuade an open-source maintainer to approve malicious code.

Criminal use

The report argues that criminal adoption of AI still trails the most advanced laboratory behaviour, but that the gap is narrowing. It points to recent cases in which attackers used commercial or older models rather than the latest frontier systems.

One example was a ransomware affiliate linked to The Gentlemen service, which Gambit Security documented using Claude Code against at least six organisations. The operator, according to the report, chose an older, less restricted model and opened a new session to assert authorisation when refusals occurred before letting the model run the intrusion.

It also highlighted JADEPUFFER, described as the first documented case of agentic ransomware. In that operation, a model carried out an extortion campaign from start to finish once a human operator had initiated it.

These cases suggest AI is taking a larger role in ransomware, shifting from a supporting tool to a system that can handle substantial parts of an intrusion. The report stops short of saying criminal groups have matched the autonomy seen in the laboratory incidents, but says the line is moving.

Access market

Another focus was the emergence of a market around AI access itself. The report describes a layered underground trade in stolen credentials, resale services and packaged offensive tools built around compromised or modified access to commercial models.

It cited an operation tracked as Zerofot that harvested almost 3,000 valid API keys and credentials across more than 1,700 hosts in about seven weeks. Those credentials were then said to feed gateway services that pool access and conceal the buyer's identity.

The report also said sellers are bundling access into finished products, including a jailbroken Claude model marketed as a penetration testing platform. It argued that this shows access, compliance and autonomy becoming distinct commodities in criminal markets.

Attack targets

Beyond misuse of models, AI systems themselves are becoming targets. Coding agents and workplace copilots are exposed because they treat files, repositories, pull requests and shared content as trusted inputs.

The report highlighted GhostApproval, a technique that uses symbolic links in a malicious repository to induce an agent to write files outside its workspace. It also noted that Google's Gemini CLI and Anthropic's Claude Code required patches for vulnerabilities that a malicious GitHub issue could trigger.

Business productivity tools featured as well. Microsoft 365 Copilot's search function could leak files from a single crafted link, while researchers also demonstrated a self-propagating prompt injection in Microsoft Word's Copilot.

Supply chain weaknesses added another layer of risk. The report pointed to more than 140 trojanised Mastra AI framework packages attributed to North Korea's Sapphire Sleet, alongside malicious LiteLLM releases that reportedly exposed credentials across about 2,500 companies.

Patch burden

The report said AI is helping researchers find vulnerabilities faster than organisations can fix them. It cited warnings from the UK National Cyber Security Centre about a coming patch wave and linked that to large patch volumes from major software suppliers including Microsoft and Oracle.

Examples included Squidbleed and a WordPress flaw found with a frontier model. Yet only about 1 per cent of AI-discovered vulnerabilities had been confirmed as exploited, the report said, suggesting the bottleneck lies more in patching and remediation than in discovery.

That dynamic is shifting security from a scarcity problem to a speed problem. As flaws become easier to identify, the advantage moves to organisations that can validate, patch and deploy fixes quickly.

Routine leakage

The report also examined ordinary workplace use of generative AI tools. It said one in every 36 prompts submitted in July, or about 2.8 per cent, carried a high risk of sensitive data leakage, while 88 per cent of organisations recorded at least one high-risk prompt during the month.

Latin America showed the highest regional rate at one in 29 prompts, according to the findings. By sector, Business Services recorded the highest level at one in 27, ahead of Healthcare, IT and Government.

The report said these figures point to exposure created by routine use rather than deliberate attacks. "One in every 36 prompts submitted to generative AI tools in July, about 2.8 percent, carried a high risk of sensitive data leakage, and 88 percent of organizations recorded at least one high-risk prompt that month," Check Point Research said.