Red Hat AI 3.5 expands AI safety and observability tools
Wed, 9th Sep 2026 (Today)
Red Hat has launched Red Hat AI 3.5, adding tools for AI safety, observability and multi-tenant infrastructure management.
The release is Red Hat's latest effort to position its AI software as a platform for companies moving from pilot projects to broader operational use.
Among the main additions is the general availability of EvalHub, designed to help organisations assess AI models before deployment. It supports safety benchmarking and auditable compliance reporting for custom models, retrieval-augmented generation systems and AI agents.
Red Hat has also introduced dashboards to give platform teams a clearer view of inference health, GPU use and model performance. Non-admin users can also see token consumption showback and monitor distributed inference workloads.
Safety focus
The release places particular emphasis on pre-deployment testing and governance. Evaluated catalogue models now include built-in Garak benchmark scores, along with safety, personally identifiable information exposure and toxicity risk scores.
More than 20 validated models have been added to the catalogue, including models from Google, Nvidia and Alibaba Cloud. A smaller set has also been marked as validated for tool-calling, relevant for agent-based applications that need models to interact with external tools and systems.
Another part of the update addresses security around AI agents. Support for the Responses API and built-in retrieval-augmented generation is now generally available, while integrated NeMo Guardrails are intended to block malicious tool calls.
Another new element, AutoRAG, links enterprise data repositories to agentic applications. It includes multilingual document support, conversational testing and contextual retrieval, along with a visual pipeline to help teams test configurations before deployment.
Shared infrastructure
Red Hat AI 3.5 also expands support for running AI workloads across shared GPU infrastructure. Fair-share GPU scheduling is intended to manage resource allocation across tenants, while priority-aware serving routes requests according to workload importance.
This is intended to protect real-time inference jobs while allowing lower-priority background workloads to use spare capacity. Controlled model rollout has also been added to manage traffic during model updates and reduce service disruption.
For customers seeking stronger tenant separation, the software now officially supports hosted control planes on OpenShift Virtualization. This gives each tenant a dedicated cluster control plane while allowing hardware to be consolidated underneath.
Support for AI workloads inside virtual machines on OpenShift Virtualization is also intended to provide isolation at the VM level on shared GPU-enabled systems. This allows infrastructure operators to manage and upgrade the environment from a single control point.
Operational visibility
Observability is another core part of the release. New dashboards provide visibility into hardware inventories, active utilisation and available GPU capacity, along with model and agent performance monitoring and MLflow visual tracing for agent workflows.
Red Hat is also adding per-user token metering, which could help organisations track internal usage and assign costs across departments or teams. This reflects a broader shift in enterprise AI projects as finance and IT teams seek clearer controls over spending and consumption.
On the performance side, Inference-Time Scaling adjusts compute use dynamically based on query complexity. CPU offloading is now generally available, and storage offloading is being offered as a developer preview to help models handle longer conversations and larger documents without extra GPU hardware.
The release also extends model serving support beyond OpenShift to third-party Kubernetes services. This is now generally available on CoreWeave CKS and Microsoft Azure, while Amazon EKS has been added as a technology preview.
Agent templates
To support AI application development, AI Hub now includes agent templates and starter kits for common uses such as code review, document processing and research workflows. These packages combine tools, frameworks and deployment configurations intended to run in sandboxed environments with security policies in place.
Joe Fernandes, Vice President and General Manager of the AI Business Unit at Red Hat, said the market focus has changed as businesses seek to run AI systems with the same discipline expected of other critical technology operations.
"The conversation has moved from getting AI into production to running it at scale as trusted enterprise infrastructure, which requires safety evidence, governed agents, cost attribution and multi-tenancy," Fernandes said. "With Red Hat AI 3.5, we are delivering the operational controls, verifiable trust and agentic foundations IT leaders need to run AI as a safe, controlled and accountable enterprise AI architecture across the hybrid cloud."