When a platform supports AI systems at OpenAI's scale, the network has to just work. It has to connect services reliably, enforce the right controls, support different kinds of compute, and give infrastructure teams enough visibility to understand what is happening. OpenAI has adopted Isovalent Networking for Kubernetes, the enterprise Cilium-based offering from Isovalent, as its default Kubernetes networking stack. It gives OpenAI a common foundation for CNI, IPAM, and L4/L7 filtering across cloud provider environments and bare metal. For an organization operating at OpenAI's pace, that consistency matters. OpenAI needed Kubernetes networking that could keep up with infrastructure growth, support security and compliance controls, and help platform teams troubleshoot network flows without treating every environment as a separate networking problem. Why Cilium And Isovalent Cilium fits that requirement because it can work as a Kubernetes CNI across different infrastructure choices. It gives OpenAI flexibility to mix different kinds of compute while reducing dependence on any single cloud provider or bare-metal provider CNI. Isovalent makes that foundation practical for production and mission critial workloads by pairing Cilium with enterprise support, upgrade planning, and expertise from the team that created and maintains the project. That matters for AI infrastructure because the compute layer is rarely static. As new capacity, workload patterns, and service requirements arrive, platform teams need the network to adapt without becoming another bespoke part of each environment. Cilium also gives OpenAI a policy model for L4/L7 filtering and node-to-node security. OpenAI uses both CiliumNetworkPolicy and CiliumClusterwideNetworkPolicy, with a shift toward primarily using cluster-wide policies and targeted namespace policies where they make sense. The goal is to avoid policy sprawl while keeping controls precise enough to matter. OpenAI moved toward a clearer model: secure egress and ingress by default, with specific exceptions handled explicitly. That is the kind of lesson other platform teams can use. Policy has to be more than a language. It has to be a model that teams can operate without creating confusion, bypasses, or a false sense of security. The Solution: Common CNI, IPAM, Policy, And Production Visibility OpenAI's default Kubernetes networking stack gives the team a foundation that combines CNI, IPAM, and L4/L7 filtering in one place. For OpenAI, the result is a networking stack that brings several capabilities together: Cilium provides a common networking layer across Kubernetes environments in multiple clouds and bare metal. Cilium supports L4/L7 filtering for targeted network controls. CiliumNetworkPolicy and CiliumClusterwideNetworkPolicy help OpenAI define controls at the right scope. Cilium has been used to demonstrate network controls for both internal and external compliance needs. Hubble helps teams debug production flows with Kubernetes context. Hubble gives teams a practical way to see what the network is doing. When teams need to understand unexpected production network flows, hubble observe can be easier than taking traditional packet captures (tcpdump). That difference matters in production. Packet captures still have their place, but platform teams need faster ways to answer service-level questions without dropping into low-level packet analysis every time. The Outcome: More Consistent Networking And Better Controls The outcome is a Kubernetes networking stack that OpenAI can use across a complex and changing infrastructure environment. Cilium gives OpenAI a common CNI foundation across multiple clouds and bare metal. It helps the team go to production with IPAM, CNI, and L4/L7 filtering capabilities. It supports network controls that OpenAI can use in compliance conversations. And with Hubble, it gives teams a more direct way to debug production flows. For OpenAI, scalability is not just about adding more infrastructure. It is about keeping the networking layer predictable as the platform expands across different compute environments, clouds, services, and teams. "Cilium allows my team to move quickly working with GPU capacity in an extremely competitive market where bringing forward launches or compute by a day or week can have profound implications on the business. Having a single kube-native networking control plane lets us easily meet compliance or security requirements in a dynamically changing environment where agentic workflows have accelerated the rate of change using the full tooling, programmability and power of the Kubernetes ecosystem." Fangyuan Li, Fleet GPU Expansions tech lead, OpenAI What Other AI Platform Teams Can Learn AI infrastructure teams often focus on compute first, and for good reason. But the network becomes part of the platform's ability to absorb growth. As services multiply, environments diversify, and production expectations rise, teams need a networking foundation that can stay consistent across change. OpenAI's experience points to several useful lessons for other platform teams: First, standardizing on a common CNI can reduce the operational differences between cloud environments and bare metal. That gives platform teams a more familiar operating model as capacity, workloads, and infrastructure choices change. Second, policy design has to be operable. A large collection of broad namespace-level policies can create complexity without necessarily improving security. OpenAI's move toward primarily using cluster-wide policies, with targeted namespace policies for specific cases, reflects a more deliberate operating model. Third, production troubleshooting needs Kubernetes-aware visibility. Hubble gives teams a way to inspect flows without reaching first for packet captures, which can make investigations faster and more accessible to platform teams. Finally, enterprise support matters when the network is part of the production foundation. Working with Isovalent means working with the creators and maintainers of Cilium, with the engineering depth and operational experience required for demanding environments. Build A More Predictable Kubernetes Network With Isovalent Kubernetes networking challenges rarely arrive in isolation. They show up through cloud and bare-metal differences, policy complexity, compliance needs, production troubleshooting, upgrade planning, and the pressure to keep platforms moving as infrastructure changes underneath them. If your platform team is facing similar challenges, Isovalent can help. Isovalent Networking for Kubernetes gives teams an enterprise Cilium-based foundation for Kubernetes networking, security, and observability, backed by Cilium expertise for production and mission-critical environments. Learn More: ● Read more Isovalent case studies ● Learn more about Isovalent Networking for Kubernetes ● Learn more about Isovalent Reliable Connectivity ● Get hands-on with Isovalent labs