SecurityBrief US - Technology news for CISOs & cybersecurity decision-makers
United States
ChatSee warns of AI failures beyond hallucinations

ChatSee warns of AI failures beyond hallucinations

Thu, 30th Jul 2026 (Today)
Mark Tarre
MARK TARRE News Chief

ChatSee has published a report on enterprise AI failures based on more than 10,000 observed incidents. It found that hallucination-related failures accounted for less than 10% of the cases analysed.

The report argues that the main sources of failure are changing as companies move from chatbot-style systems to AI agents that carry out tasks. Execution and action failures rose 62% relative to ChatSee's Q2 2024 baseline, while resolution and escalation breakdowns made up 31.1% of observed failures.

The findings draw on incidents collected between 2023 and 2026 across 10 industry sectors, business functions, and seven stages of the AI lifecycle. ChatSee used public incident reports, community-reported failures, open-source agent traces, datasets, AI interaction logs, production observations, and its own research to build a taxonomy of more than 150 failure categories.

Many organisations still assess AI risk mainly through hallucinations, output controls, and prompt testing. ChatSee's analysis suggests that this emphasis misses a broader set of problems that emerge when systems are asked not just to generate answers but to complete work inside business processes.

Among the patterns identified, financial services showed failures concentrated around escalation and governance. In healthcare and insurance, more failures were tied to context integrity and policy-sensitive guidance, while technology and telecom were more exposed to execution, entitlement, and workflow completion problems.

The report also found that the same business function can fail differently depending on the industry. Customer support systems in banking, telecoms, travel, healthcare, and logistics may share similar AI architectures, but differences in data, policy, and escalation structures create different risks.

Beyond hallucinations

ChatSee said the changing distribution of failures reflects a wider shift in enterprise AI use. A system may produce a fluent and compliant response yet still fail to resolve a customer issue, trigger the wrong tool, miss a required escalation, or leave a workflow unfinished.

That distinction matters because many current safeguards were designed for earlier chatbot deployments. Output filters and red-teaming remain relevant, but the report argues they do not address all the risks that emerge once software agents are connected to tools, data sources, and internal workflows.

Sekhar Sarukkai, Chief Executive Officer and Co-Founder of ChatSee, said the issue now extends beyond answer quality. "The enterprise AI question is no longer only whether a model can answer correctly, but whether the AI system can retrieve the right context, take the right action, escalate at the right time, and bring work to resolution," he said.

He gave an example from banking in which the customer experience may appear satisfactory even when the required outcome never happens. "A banking customer can report suspicious activity, receive a polite and compliant response, and still never be escalated for human review. That is a serious enterprise failure even if the model never hallucinated," Sarukkai said.

Industry patterns

The study says enterprises need to assess failures in the context of specific workflows rather than through a single generic model-risk lens. It argues that AI systems operating in regulated sectors or handling customer service need monitoring that captures how they behave once deployed, including whether they complete tasks, escalate exceptions, and follow business rules.

ChatSee linked that argument to a recent incident involving OpenAI and Hugging Face, which it said showed how autonomous systems can create risk through the way they pursue an assigned objective. Governance models based on static checks are likely to be insufficient if agents are given greater autonomy over tools, data access, and decision paths.

Sarukkai said the point is becoming more pressing as companies give AI systems more discretion. "As the recent OpenAI-Hugging Face incident showed, agents do not need malicious intent to create enterprise risk," he said.

He added: "An AI system can pursue an assigned goal through a strategy the operator never intended. That is the failure class enterprises will increasingly face as agents gain more autonomy. Governance cannot be a static checklist; it has to account for how systems behave while they act."