SecurityBrief US - Technology news for CISOs & cybersecurity decision-makers
United States
Google expands AI infrastructure with Lustre & C4N

Google expands AI infrastructure with Lustre & C4N

Sun, 2nd Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

Google Cloud has moved its Managed Lustre storage service into general availability and has also made its C4N network- and storage-optimised virtual machines generally available. The updates are part of a broader monthly round-up of additions to Google's AI infrastructure portfolio.

Managed Lustre is offered in four performance tiers, ranging from 125 MB/s to 1000 MB/s per TiB of capacity, and can scale to 8 PB of storage. The service uses DDN's EXAScaler technology.

C4N is Google's first virtual machine series designed specifically for network- and block storage-heavy workloads. The machines use 5th Gen Intel Xeon Scalable processors and Google's Titanium hardware. Google cited 400 Gbps network bandwidth, 95 million packets per second and up to 25 GiB/s of block storage throughput when used with Hyperdisk Extreme.

Google also raised the upper size limit for standard Google Kubernetes Engine clusters using Dataplane V2 with network policies enabled. The service now supports clusters of up to 15,000 nodes, a scale aimed at large corporate, AI and machine learning deployments.

Another change focused on the use of expensive accelerator hardware. A co-operative time-slicing feature in llm-d allows reinforcement learning workloads to interleave separate jobs on shared physical systems, raising accelerator duty cycles from about 40% to 70% without affecting model convergence or accuracy.

Security and tooling

Google has also open-sourced k8s-aibom, a Kubernetes controller intended to identify AI runtimes running in container clusters and create CycloneDX machine learning bills of materials. The tool is intended to help organisations track AI software components inside Google Kubernetes Engine environments.

Google also highlighted several developer and operations guides tied to its TPU and Kubernetes services. These include support materials for Moonshot AI's Kimi K3 open-weight model, guidance on running Ray on TPUs, a microbenchmark suite for TPU performance testing, and instructions for using GKE Agent Sandbox and Pod snapshots to increase the number of AI agents that can run on a fixed compute footprint.

The update also pointed to a technical blueprint describing inference work for Mistral 3 Large on Ironwood, also known as TPU v7x. According to Google, engineering changes including hybrid sharding, tree reductions and asynchronous scheduling produced a 1.5x performance gain and increased throughput by up to 48% while maintaining benchmark accuracy neutrality.

Market positioning

Alongside the product releases, Google drew attention to external and internal research on AI infrastructure demand. It said Gartner had named it a Leader in its inaugural Magic Quadrant for AI Infrastructure and placed it highest for Ability to Execute and furthest for Completeness of Vision.

Google also cited findings from a survey of more than 1,400 senior IT leaders in its State of AI Infrastructure report. According to the survey, 83% of organisations said they needed infrastructure upgrades to support production-grade agentic AI.

Some of the additions highlighted from the previous month focused on confidential computing and observability. Confidential Computing is now available on G4 accelerator-optimised machine series with NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, while a new OpenTelemetry-based TPU AI Telemetry Collector Agent can send TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus or self-hosted Grafana deployments.

Google also promoted a new TPU Developer Hub as a central resource for model builders and developers working with its tensor processing units.

Inference and orchestration

Several earlier updates centred on inference and orchestration for AI workloads. GKE Agent Sandbox is now generally available, while Google AI Edge Portal now supports benchmarking and debugging of on-device large language models.

Google also pointed to Cloud Storage Rapid, a family of storage products for AI workloads that includes Rapid Bucket, a zonal object storage service, and Rapid Cache, which is designed to speed up reads and place compute closer to data in existing buckets.

On infrastructure architecture, Google highlighted internal analysis from senior engineering leaders on adapting data centre fabrics, wide area networking and global network design to AI workloads. It also referred to a cluster-level reliability model for training frontier models on TPUs rather than relying on instance-level reliability.

Benchmarking claims were another part of the update. Google cited an independent benchmark report that found GKE Inference Gateway delivered 15.7% higher throughput, 92.8% shorter wait times and 62.6% lower inter-token latency than the next-leading managed Kubernetes service, which it attributed to the use of prefix caching for large language model inference.

Customer references included Pager Health, Trustpilot and Imgix, which Google said are using combinations of GKE, BigQuery, Cloud SQL, Gemini Enterprise Agent Platform, Dataflow and G4 virtual machines for workloads ranging from healthcare services to large-scale media processing. One example cited by Google said Imgix serves more than 8 billion images and videos a day using AI Hypercomputer systems with G4 virtual machines equipped with NVIDIA RTX PRO 6000 Blackwell GPUs.