How to secure Ray and vLLM on Kubernetes

Distributed inference opens a second network that four default scanner configurations did not report. Here is how to find it and close it.

Published · Based on our Hidden Network research

Last reviewed: · 5 min read

Sorami’s Hidden Network study found that a pod in an unrelated namespace could reach vendor-default Ray control ports in one EKS cluster, speaking cleartext gRPC. To close them, apply an ingress default-deny NetworkPolicy first. Then check Ray ports from inside the pod, encrypt traffic between nodes, and upgrade vLLM before trusting its API key.

The measurements come from our research report on Amazon EKS, with raw logs in the public data repo. Each step says what it closed.

Key takeaways

  • Ray GCS and raylet ports were reachable from an unrelated namespace, in cleartext.
  • An ingress default-deny NetworkPolicy blocked every Ray port that was open.
  • Default scanners did not inspect the RayCluster, so its Ray pods were not analysed.
  • Cilium plus WireGuard carried streaming, with a 3% to 8% difference in one run.
  • A vLLM API key returning 401 does not show that the API is protected.

What did the study find?

The study deployed KubeRay 1.7.1 and Ray 2.52.0 with vLLM using defaults. From an unprivileged pod in an unrelated namespace, Ray GCS on port 6379 and raylet RPC on ports 10002 to 10006 answered in cleartext.

Four default scanner configurations did not inspect the RayCluster, so its Ray pods were not represented in their analysis. Custom rules were not tested. Of 17 Ray listening sockets, 15 were missing from the declared containerPort list, which is metadata and not a security boundary. The Ray security documentation says Ray should only run inside a trusted network.

How do you close the Ray and vLLM network?

Work in this order.

1. Apply an ingress default-deny NetworkPolicy

Deny all ingress in the Ray namespace, then add two allows. Ray pods may reach each other on every port, because Ray picks ports at runtime. The vLLM port is open only to named callers such as the router.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: inference
spec:
  podSelector: {}
  policyTypes: ["Ingress"]

In our test it blocked every Ray port that was open. It showed no measurable steady-state difference in one run per level. It is ingress only, so add an egress policy if you need containment. The policies are in the test harness. On EKS, consider NETWORK_POLICY_ENFORCING_MODE=strict so new pods start default-deny. We did not test the startup window of newly created pods.

2. Lock down the Ray Job API

Treat an exposed unauthenticated Ray Job API on 8265 as a potential arbitrary-code-execution interface. This rests on documented API semantics; no job was submitted. The KubeRay default image had it; the vLLM image did not. Authenticate it or disable the dashboard. Evaluate Ray token authentication (RAY_AUTH_MODE=token), which is off by default. We did not test it.

3. Find ports from the running pod

List listening sockets inside each Ray pod, or scan from a pod in another namespace against your own targets.

4. Encrypt traffic between nodes

The tested Linkerd sidecar configuration returned HTTP 503 for every request. That is an engineering incident with no isolated root cause, not a security result. Cilium WireGuard chained behind the VPC CNI (the combined configuration, not WireGuard alone) carried every streaming request. In one bracketed run it showed about 3% to 8% lower throughput. When removing Cilium, set cni.uninstall=true. The default left a stale CNI config that broke new pods.

5. Upgrade vLLM, then set an API key

Set --api-key from a Kubernetes Secret on vLLM 0.22.0 or later. Versions from 0.3.0 up to 0.22.0 are affected by CVE-2026-48746, a Host header bypass. Treat this as a supporting layer. In our test the normal key path returned 401 on two /v1 routes. That is not full authentication. Test every route, including /metrics and /tokenize, with no key and a wrong key.

Sorami’s view

This section is opinion. It is how we would order the work for a team running Ray or vLLM today.

Treat the Ray network as plumbing that no other workload should reach. The NetworkPolicy is plain Kubernetes and takes an afternoon, so do it first. An API key sent in cleartext on a flat network is a weak layer on its own. A reachability test from a neighbour pod is part of our AI production readiness review and our cloud penetration testing.

Checklist

  • Apply an ingress default-deny NetworkPolicy to every Ray namespace.
  • Allow only named callers to reach the vLLM API port.
  • Disable or authenticate the Ray Job API on port 8265.
  • List listening sockets inside each Ray pod.
  • Encrypt inter-node pod traffic with WireGuard or IPsec.
  • Upgrade vLLM to 0.22.0 or later before setting an API key.
  • Test every vLLM route with no key and a wrong key.

Related reading

The full Hidden Network report has the method, tables and limits. Our AI on Kubernetes Helm chart research covers single-chart defaults. For agents, see how an agent used DNS to get around its sandbox. More is in the guides index.

Sources

Questions before you book

Practical answers.

How do I secure a Ray cluster on Kubernetes?

Start with an ingress default-deny NetworkPolicy in the Ray namespace. Allow the Ray pods to reach each other on all ports and allow the serving port only from named callers. In Sorami’s test that blocked every Ray port that was open to a pod in another namespace. Then evaluate Ray token authentication and encrypt traffic between nodes.

Which Ray ports need to be protected?

In our test the exposed ports were Ray GCS on 6379 and raylet RPC on 10002 to 10006. Treat an exposed unauthenticated Ray Job API on 8265 as a potential arbitrary-code-execution interface. That rests on the documented API semantics; no job was submitted in the study. Ray also opens ports at runtime, so list sockets inside the pod rather than trusting the manifest.

Is the vLLM --api-key flag enough to secure the API?

Not on its own. In our test the normal key path returned 401 on two /v1 routes, which does not show that the API is protected. Other routes were not tested. The tested vLLM version is in the affected range of CVE-2026-48746, a Host header bypass fixed in 0.22.0. Upgrade first, then test every route with no key and a wrong key.

Does Linkerd work with vLLM streaming?

In our test the Linkerd sidecar configuration returned HTTP 503 for every request. The root cause was not isolated and the Linkerd version was not recorded, so this is one failed configuration and not a general result. Cilium chaining plus WireGuard below the application carried every streaming request.

How much did Cilium with WireGuard change distributed inference speed?

In one same-session bracket, Cilium chaining plus WireGuard combined showed 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median latency. The difference rose with concurrency. It does not isolate WireGuard, and it comes from a single run per level on a small model.

About Sorami

Sorami is an Australian cyber security and cloud consultancy. We build and secure cloud environments, Kubernetes and AI agents, and the senior engineers who scope the work deliver it themselves. To discuss this research or a review of your own stack, use the form below. Only your email is required, and we reply within one business day.

An enquiry, not a booking. We use your details only to reply. See our privacy notice, or go to the contact page.

Let’s scope it

Want a neighbour-pod test of your inference cluster?

Tell us how your Ray or vLLM cluster is laid out. We will scope a review of what another workload can reach.

Request a quote