The Hidden Network: Ray Control-Plane Exposure in Distributed LLM Inference on Kubernetes
What a pod in an unrelated namespace could reach on default Ray and vLLM deployments in one Amazon EKS testbed, what four default scanner configurations reported, and what an ingress NetworkPolicy blocked.
In one Amazon EKS cluster, a pod in an unrelated namespace could reach default Ray control-plane listeners, which spoke cleartext gRPC. An ingress NetworkPolicy blocked those ports in steady state. Four default scanner configurations did not inspect the RayCluster resource. Sorami measured this once, on one topology.
Key takeaways
- Ray GCS and raylet ports answered cleartext gRPC across namespaces.
- An ingress NetworkPolicy blocked every Ray port that was open.
- Default scanners did not inspect the RayCluster, so its Ray pods were not analysed.
- 15 of 17 Ray sockets were missing from declared ports, which are only metadata.
- Cilium chaining plus WireGuard differed by 3% to 8% in one run.
- Job API code execution was not shown against the GPU inference stack.
Want the steps without the method? Our practical guide, how to secure Ray and vLLM on Kubernetes, turns these findings into a checklist. The full report follows.
Summary
Production LLM serving is moving from single-GPU engines to multi-node topologies built on Ray, vLLM, and KubeRay. In the tested KubeRay and vLLM topology, distributed inference introduced Ray control-plane and tensor-transfer listeners that were not represented by the workload's declared port metadata. We measured this surface on one dedicated Amazon EKS 1.35 cluster running vendor-default charts and images, from one constructed neighbour pod in an unrelated namespace.
Under these defaults, in this cluster, topology and version set, that pod could reach Ray control-plane listeners: a two-node Ray pipeline-parallel vLLM deployment exposed Ray GCS (6379) and raylet RPC (10002 to 10006) cross-namespace, and they spoke cleartext gRPC. An ingress-only default-deny NetworkPolicy (L1) turned every probed Ray port that was open at L0 into BLOCKED from the neighbour pod in steady state; the pod startup window was not tested.
The Ray Job API (8265) answered unauthenticated requests on the default CPU Ray image but was not listening in the GPU vLLM deployment, so remote code execution against the GPU inference stack was not demonstrated (claim table in The distributed control plane: cleartext and reachable). Four default scanner configurations (Trivy, Checkov, Kubescape, kube-linter) did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis, and they did not identify the RayCluster's runtime network exposure. These are runtime and network properties, and custom rules were not tested (Four default scanner configurations and the runtime surface). 15 of 17 observed Ray listening sockets were absent from the pods' declared port metadata: visibility drift for tooling that treats declared ports as an exposure inventory (containerPort is metadata, not a security boundary). A single-GPU vLLM engine exposed 26 routes with no authentication configured, including tokenizer and prefix-cache telemetry.
Two hardening results are weaker. An engine API key (L3) turned ordinary unauthenticated requests on two /v1 routes (/v1/models, /v1/completions) from 200 into 401, which shows the normal key path works and does not show that API-key authentication protects vLLM: the other routes were not tested, and the tested engine, vllm/vllm-openai:v0.11.0, is in the affected range of CVE-2026-48746 (a Host-header API-key bypass fixed in 0.22.0), which we did not test. The layers L1 and L3 showed no measurable steady-state difference within cross-session variance (n=1 run per level, median full-request (E2E) latency within 5% of baseline). Cilium chained on the VPC CNI with WireGuard (L2b, the combined configuration, not WireGuard alone) served 512 of 512 streaming requests, and interface counters showed inter-node pod traffic routed through the WireGuard device. NCCL traffic was not separately identified or captured, so NCCL encryption is inferred from that datapath and not measured. Against a same-session L0 bracket on a different GPU pair, the encrypted configuration produced 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median full-request (E2E) latency in this single bracketed run, rising with concurrency. The low-concurrency part of that difference is within L0-to-L0 drift, and the closing L0 bracket was not a proven Cilium-free state.
Separately, an engineering incident: the tested Linkerd mTLS sidecar configuration (L2a) failed to carry the vLLM workload (128 of 128 requests failed with HTTP 503 at every concurrency level), the root cause was not isolated, and it is not a security finding. An incomplete Cilium removal procedure (uninstall with the documented chart default cni.uninstall=false) left a stale CNI config on every node and broke new pod sandboxes for about 28 minutes. Results are limited by single runs, no persisted per-request samples, heterogeneous GPUs, a small model, and a reused baseline for L1 and L3, all disclosed.
What this study shows and does not show
Scope, environment and evaluation bounds:
- Threat model: This study evaluates what an attacker who controls one ordinary pod in an unrelated namespace of the same cluster can reach. Container escape, node compromise, attacks on the Kubernetes API server, cloud IAM, supply chain, model-level attacks and attacks from outside the VPC are out of scope.
- Environment: One dedicated Amazon EKS 1.35 cluster in ap-southeast-2, one day (2026-09-28 UTC), on-demand GPU nodes for the distributed runs, a small model (Qwen2.5-1.5B in fp16) and short prompts. Results say nothing about spot interruption, larger models or other regions.
- Evaluation bounds: Each layer and concurrency level has one 128-request run, with no repeated trials, confidence intervals or significance tests. The Stage D L0 baseline is a reused copy of the Stage C run, so the L1 and L3 deltas are cross-session.
- What the hardening results do not show: The API key result (L3) shows that the normal key path returned 401. It does not show that API-key authentication protects vLLM: only two
/v1routes were tested, and the tested engine is in the affected range of CVE-2026-48746, which we did not test. The encryption result is Cilium chaining plus WireGuard combined, not WireGuard alone, and it is one bracketed run. Ray token authentication and TLS were not evaluated.- What was and was not demonstrated: Ray control ports were reachable across namespaces, and the Ray Job API was reachable on the default CPU Ray image. The Job API was not listening in the GPU vLLM deployment, and remote code execution against the GPU inference stack was not demonstrated. The distributed control plane: cleartext and reachable has a claim-by-claim table.
- Scanners: The result is about four default scanner configurations. Custom rules, admission-time policy engines and other scanners were not tested.
Demonstrated, not demonstrated and engineering incidents
Demonstrated in this testbed:
- An ordinary pod in an unrelated namespace could reach Ray GCS (6379) and raylet RPC (10002 to 10006) on the GPU pipeline, and those listeners spoke cleartext gRPC (The distributed control plane: cleartext and reachable).
- The Ray Job API on 8265 answered unauthenticated requests from that pod on the default CPU Ray image (Exposed unauthenticated endpoints).
- 15 of 17 observed Ray listening sockets were absent from the declared port metadata (Four default scanner configurations and the runtime surface).
- Four default scanner configurations did not inspect the RayCluster custom resource, so the generated Ray pods were not represented in their analysis (Four default scanner configurations and the runtime surface).
- An ingress default-deny NetworkPolicy blocked every probed Ray port that was open at L0, in steady state (Reachability matrix).
- Cilium chaining plus WireGuard combined carried 512 of 512 streaming requests, and interface counters showed inter-node pod traffic on the WireGuard device (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
Not demonstrated:
- Remote code execution on any Ray node. The Job API was not listening in the GPU vLLM deployment, and no job was submitted to the CPU deployment (The distributed control plane: cleartext and reachable).
- That a vLLM API key protects the server. Only two
/v1routes were checked, and CVE-2026-48746 was not tested (Security effect of each layer). - NCCL encryption, WireGuard alone, the pod startup window, egress containment, and Ray token authentication or TLS (L2b: Cilium chaining plus WireGuard combined (single bracketed run), Limitations and Future work).
Engineering incidents, not security findings:
- The tested Linkerd sidecar configuration failed to carry the vLLM workload, with HTTP 503 for every request; the root cause was not isolated (L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)).
- An incomplete Cilium removal left a stale CNI config on every node and broke new pod sandboxes for about 28 minutes (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
Introduction
Inference engines such as vLLM expose an OpenAI-compatible HTTP API and, when scaled past one GPU, delegate placement and coordination to Ray. The Kubernetes operator KubeRay packages Ray as a custom resource. In the deployment we tested there were two sets of listeners: the declared one (a Service on port 8000) and a second set that the manifests did not declare (Ray GCS, raylet, object manager, worker RPC, and the collective-communication channel that carries activations or KV cache between GPUs). Ray's own documentation states that a Ray cluster should run only inside a trusted network. Kubernetes allows pod-to-pod traffic unless a NetworkPolicy restricts it. This study measures what that combination allowed in one testbed.
This report asks three questions:
- RQ1: What does an unprivileged neighbour pod see when single-node and multi-node LLM inference stacks are deployed with vendor defaults?
- RQ2: Do four widely used static Kubernetes scanners, in their default configurations, identify that runtime network exposure?
- RQ3: What does each hardening layer close, and what difference in throughput and latency did we observe on the same topology?
Contributions:
- An empirical attack-surface inventory of vendor-default KubeRay 1.7.1, Ray 2.52.0, vLLM production-stack 0.1.12, and
vllm/vllm-openai:v0.11.0on EKS, from an unprivileged cross-namespace pod (Results: attack surface). - Evidence that four default scanner configurations did not inspect the RayCluster custom resource, so the generated Ray pods were not represented in their analysis and the RayCluster's runtime network exposure was not identified (no scanner output contains a reference to the RayCluster resource), and that 15 of 17 observed Ray listening sockets (10 of 11 on the head, 5 of 6 on the worker) were absent from the pod's declared port metadata (
containerPort). This is visibility drift for tooling that treats declared ports as an exposure inventory.containerPortis metadata, not a security boundary, and Kubernetes does not block undeclared ports (Four default scanner configurations and the runtime surface). - Wire-level evidence that the multi-node Ray control plane is cleartext gRPC and reachable cross-namespace, and a claim-by-claim statement of what was and was not shown about the Ray Job API and remote code execution: exposed on the default CPU Ray image, not listening in the GPU vLLM deployment, and no code execution demonstrated against the GPU inference stack (The distributed control plane: cleartext and reachable).
- A layered hardening evaluation (L0 to L3) on a fixed two-node pipeline, with a per-layer security observation scoped to what was probed, and a performance comparison that is explicit about its statistical limits (Results: hardening). The L3 API key shows that the normal key path returned 401 on two
/v1routes; it does not show that API-key authentication protects vLLM, and the tested vLLM version was publicly known to be bypassable at experiment time (CVE-2026-48746, not tested here). - A same-session, bracketed comparison of Cilium chained on the EKS VPC CNI with WireGuard (the combined configuration, not WireGuard alone) on Ray pipeline-parallel inference: it carried SSE streaming, and in this single bracketed run it produced about 3% to 8% lower throughput and higher median E2E latency, with the drift, the Cilium-plus-WireGuard confound, and the unproven Cilium-free state of the closing bracket stated (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
- An operational lifecycle hazard for Cilium chaining on EKS: an incomplete removal procedure (
helm uninstallwith the documented chart defaultcni.uninstall=false) leaves a CNI config that breaks every new pod sandbox cluster-wide until removed (L2b: Cilium chaining plus WireGuard combined (single bracketed run)). - An engineering incident, not a security finding: the tested Linkerd configuration failed to carry the vLLM workload (every request returned 503); the root cause was not isolated, and the Linkerd version and proxy logs were not kept (L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)).
Background
Distributed LLM inference
vLLM serves an OpenAI-compatible API over HTTP on port 8000 by default, with streaming responses delivered as Server-Sent Events (SSE). Beyond completions it exposes tokenizer, embedding, scoring, and operational routes on the same port, plus a Prometheus /metrics endpoint. Authentication is optional and off by default; it is enabled with --api-key or VLLM_API_KEY, which requires a bearer token on the /v1 OpenAI routes. The check is not a blanket control for every route on the server. CVE-2026-48746 (GHSA-94f4-hr76-p5j6, CVSS 9.1, critical, published 2026-06-02) is an API-key bypass in the OpenAI API server: a crafted Host header makes AuthenticationMiddleware reconstruct a path that skips the /v1 API-key check. Affected versions are >=0.3.0 and <0.22.0; it is fixed in 0.22.0. Deployments behind an RFC-conforming reverse proxy such as nginx are not affected [19] [C].
When a model is split across GPUs on different nodes, vLLM can use Ray as its distributed executor. The Ray components relevant here are:
| Component | Default port | Role |
|---|---|---|
| GCS (Global Control Store) | 6379 | Cluster metadata, node and actor registry, gRPC |
| Dashboard and Job Submission API | 8265 | HTTP UI and REST API to submit arbitrary jobs |
| Ray Client server | 10001 | Remote driver connections |
| Raylet, object manager, worker RPC | 10002 and up, plus ephemeral ports | Task scheduling, object transfer |
| Metrics exporter | 8080 | Prometheus |
The Job Submission API accepts an entrypoint shell command and runs it on the cluster. Without authentication in front of it, any client that can reach 8265 can run code on Ray nodes. Pipeline parallelism (PP) places consecutive layer groups on different GPUs and passes activations between stages. Tensor parallelism across nodes all-reduces every layer and is far more latency sensitive. Disaggregated prefill and decode separates the two phases onto different workers and ships KV cache between them through a connector (in vLLM v0.11.0, P2pNcclConnector, built on NCCL plus a ZMQ side channel). NCCL has no built-in authentication or encryption over TCP sockets.
Kubernetes networking and NetworkPolicy
Kubernetes requires that every pod can reach every other pod without NAT. Isolation is opt-in through NetworkPolicy objects, which are enforced only if the CNI plugin implements them. On EKS, the VPC CNI enforces policy through the aws-eks-nodeagent (eBPF) when enableNetworkPolicy=true. A pod is isolated for ingress only once some policy selects it with policyTypes: Ingress, and for egress only once some policy selects it with policyTypes: Egress; the two directions are independent. NetworkPolicy operates on L3/L4 (IP, port, protocol, namespace and pod selectors); it provides no authentication of the caller and no encryption. A containerPort entry in a pod spec is informational metadata: omitting it does not stop a process from being reachable on that port.
The VPC CNI has a NETWORK_POLICY_ENFORCING_MODE setting. In the default standard mode, a new pod starts default-allow and policies are programmed in parallel with pod startup; in strict mode a new pod starts default-deny until its policies are programmed [18] [C]. This matters for any claim about a newly created pod, because in standard mode there can be a startup window during which a pod is reachable despite a matching deny policy.
Ray has optional token authentication from Ray 2.52.0 (RAY_AUTH_MODE=token). It is disabled by default and, when enabled, authenticates the dashboard, GCS and other control-plane services; KubeRay supports configuring it [4] [C]. Ray also supports TLS for its gRPC channels. Both are native alternatives or complements to network-layer controls.
Encryption in transit between pods can be added at the node level (WireGuard in Cilium or Calico, encrypting all node-to-node pod traffic) or at the workload level by a service mesh sidecar (Linkerd, Istio) that terminates mTLS and, for HTTP, proxies at L7.
Threat Model
Attacker. The attacker controls one ordinary pod in the same cluster, in a namespace unrelated to the inference stack. This models a compromised application container, a malicious dependency in a sidecar, or a tenant workload in a shared cluster. The pod runs as uid 1000, with every Linux capability dropped (CapBnd 0000000000000000) and no mounted ServiceAccount token (01-identity, probe-identity). The attacker has no Kubernetes API access, no node access, and no cloud credentials. The attacker can open TCP connections to pod IPs and Service names.
Assets.
| Asset | Why it matters |
|---|---|
| Inference capacity | Free GPU compute, and denial of service to legitimate callers |
| Prompt and completion content | May contain regulated data; exposed via the API, via tokenizer routes, and on the wire |
| Workload telemetry | Request volume, token counts, and prefix-cache hit counters leak activity and may enable a cache side channel |
| Distributed control plane | Ray GCS and the Job API govern what code runs on GPU nodes |
| Inter-node tensors | Activations or KV cache derived from prompts, carried between GPUs |
Out of scope. Container escape and node compromise; attacks on the Kubernetes API server, etcd, or cloud IAM; supply-chain attacks on images or model weights; model-level attacks (prompt injection, jailbreaks, extraction); side-channel exploitation (the prefix-cache channel is identified but not tested); attacks from outside the VPC; and any system not created for this study.
Method
Environment
| Item | Value | Evidence |
|---|---|---|
| Cluster | Dedicated EKS cluster sorami-lab, Kubernetes 1.35 (v1.35.8-eks), Amazon Linux 2023, containerd 2.2.7 | 04-gpu-nodes |
| Network | Dedicated VPC 10.99.0.0/16, region ap-southeast-2 (Sydney), no peering with other VPCs | network.tf |
| CNI | Amazon VPC CNI with aws-eks-nodeagent --enable-network-policy=true. The addon exposes NETWORK_POLICY_ENFORCING_MODE, but its value was not recorded, so whether pods started default-allow (standard) or default-deny (strict) is unknown | 15-netpol-agent, 01-vpccni-encryption-options (variable name only) |
| CPU pool | 2 spot m5/m6i-class xlarge nodes (Ray CPU cluster, router, probe, scanners) | 00-nodes |
| Stage B GPU | 1x g6.xlarge spot, NVIDIA L4 24 GB | 00-target |
| Stage C/D GPU | g5.xlarge on-demand (A10G 24 GB, PP stage 0, Ray head) plus g4dn.xlarge on-demand (T4 16 GB, PP stage 1, Ray worker) | 04-gpu-nodes |
| L2b session GPU | 2 x g5.xlarge on-demand (A10G + A10G), head and worker pinned to the same two nodes for all three runs | 02-decision, 00-target.log in each bench directory |
Spot GPU capacity in Sydney was scarce during the study: the spot placement score was 1 of 10 for every g4dn, g5, and g6 size in every AZ, and a second spot GPU node never launched (spot-placement-scores). Stage C/D therefore ran on two on-demand nodes from a fallback node group, which is why the pair is heterogeneous.
Software under test
| Component | Version | Configuration |
|---|---|---|
KubeRay operator and ray-cluster chart | 1.7.1 | chart defaults (Stage A) |
| Ray | 2.52.0 | reported by /api/version. Token authentication (RAY_AUTH_MODE=token) is available in this version and was left at its default (off); Ray TLS was also off. Neither was evaluated |
| vLLM production-stack chart (router) | 0.1.12 | chart defaults |
| vLLM engine image | vllm/vllm-openai:v0.11.0 (tag only, digest not recorded; inside the CVE-2026-48746 affected range and exposed directly on its Service with no reverse proxy) | Stage B: Qwen/Qwen2.5-7B-Instruct, bf16, --max-model-len 8192, --gpu-memory-utilization 0.90. Stage C/D: Qwen/Qwen2.5-1.5B-Instruct, --dtype half, --pipeline-parallel-size 2, --distributed-executor-backend ray, --enforce-eager, --gpu-memory-utilization 0.6, VLLM_ATTENTION_BACKEND=TORCH_SDPA, VLLM_USE_FLASHINFER_SAMPLER=0 |
| Scanners | Trivy 0.67.0, Kubescape v3.0.40, kube-linter 0.8.3, Checkov (version not logged) | in-cluster Jobs on the rendered manifests |
The Stage C/D engine settings are workarounds for the T4 (compute capability 7.5): the FlashInfer sampler crashed on its architecture check and the xformers backend raised NotImplementedError under PP; both are documented in stageC-distributed. The 1.5B model and fp16 were forced by the 16 GB T4 (no bf16 support).
Topology selection
The preferred Stage C topology was disaggregated prefill and decode with KV-cache transfer through P2pNcclConnector. It failed on both GPUs: the connector's pinned staging pool (tensor_memory_pool._allocate_pinned_memory) raised CUDA error: out of memory at kv_buffer_size values from 2e9 down to 1e8 and GPU memory utilisation from 0.85 down to 0.4 (06-p2p-oom-evidence, 07-topology-decision). We fell back to Ray-backed PP=2 across the two nodes, one GPU each. Tensor parallelism across nodes was not attempted because each node has one GPU and cross-node TP over TCP without EFA would measure the NIC rather than the security layers. Pod placement was pinned by kubernetes.io/hostname and was identical for every Stage D layer (same two node names in every 00-target.log).
Benchmark harness
The client is bench/bench.py in the test-harness/ folder of this repository. It is stdlib-only Python so the same bytes run in python:3.12-slim for every layer, delivered through a ConfigMap. It runs as a Kubernetes Job in the neighbour namespace, so the measured path is the same cross-namespace pod-to-Service path the attacker uses.
| Parameter | Value |
|---|---|
| Endpoint | streaming POST /v1/completions on the engine Service, port 8000 |
| Prompt | fixed English sentence repeated 8 times (about 89 prompt tokens) |
| Output | max_tokens=128, ignore_eos=true, temperature=0 |
| Token accounting | server-reported usage.completion_tokens via stream_options.include_usage |
| Warmup | 8 requests at concurrency 1, discarded |
| Sweep | concurrency 1, 4, 16, 32; 128 requests per level; closed loop (each worker thread issues its next request when the previous one finishes) |
| Timeout | 300 s per request |
Metrics per request: TTFT (time to the first SSE chunk carrying text), E2E (request start to stream end), and TPOT = (E2E minus TTFT) / (tokens minus 1). Per level the harness reports p50, p95, p99, and mean of each, plus output tokens per second (sum of completion tokens divided by wall time) and requests per second. Percentiles are nearest-rank over 128 samples, so p99 is effectively the second-largest value and is sensitive to one or two outliers.
Raw per-request samples were never persisted. Inside the pod, bench/bench.py line 119 writes both the summary and the per-request list (json.dump({"summary": summary, "raw": results}, f)), but the Job command in manifests/stageB-bench-job.yaml line 49 prints only the ['summary'] object to the pod log, and the runner scripts split result-c*.json out of those RESULT log lines. The pod filesystem was discarded with the Job. Every result-c*.json in this study is therefore a 4-percentile-plus-mean summary. No latency distributions, confidence intervals, bootstrap statistics, or outlier analysis can be computed from the archived data, and none are reported. Both paths are in the local harness repo, which has no git remote and no commits, so they are cited as local paths only.
GPU utilisation was sampled with nvidia-smi inside the engine pods every ~20 s during each run. /metrics was scraped before, during, and after.
Hardening layers
One variable was changed per layer, cumulatively where stated.
| Layer | Change | Artifact |
|---|---|---|
| L0 | None (distributed baseline) | manifests/stageC-ray-pp.yaml |
| L1 | Default-deny ingress in the engine namespace (ingress only; egress left open); allow engine pods to each other on all ports; allow port 8000 only from the neighbour and engine namespaces | manifests/stageD-L1-networkpolicy.yaml |
| L2 | Linkerd mTLS sidecars on both engine pods and the bench client; Ray ports 6379 and 10002 to 10010 skipped by the proxy | Linkerd edge control plane, MESH_BENCH=1 in scripts/stageC-bench-run.sh |
| L2b | Cilium 1.20.2 chained on aws-cni with WireGuard pod-to-pod encryption; L1 removed and no API key for this run; different GPU pair, own L0 bracket | manifests/stageD-L2b-cilium-values.yaml, manifests/stageD-L2b-ray-pp.yaml (test-harness) |
| L3 | L1 plus engine --api-key from a Secret; client sends the bearer token | VLLM_API_KEY secretKeyRef in manifests/stageC-ray-pp.yaml, BENCH_AUTH=1 |
The L1 allow of all ports between engine pods is deliberate: Ray and NCCL use ephemeral ports chosen at runtime, so a port list would break the pipeline. It also means L1 does not restrict traffic inside the trust boundary of the engine itself.
L1 is ingress-only. Every policy in manifests/stageD-L1-networkpolicy.yaml sets policyTypes: ["Ingress"] (the default-deny policy at line 22), and the comment at line 11 records that egress was deliberately left open so the engine could pull the model and resolve DNS. L1 therefore tests who can reach the engine. It does not show containment of a compromised engine, which would need egress policy and an egress test; neither was done.
Baseline provenance (important)
The L0 control for Stage D is not a fresh run. The four result-c*.json files and 10-table.md in stageD-00-L0 are byte-identical to those in stageC-06-bench. We verified this with SHA-256:
| File | SHA-256 (both directories) |
|---|---|
| result-c1.json | 2157d70330b94ee1c7975c452eaddcc8f676fd7905609c96c1edb21467ccfedf |
| result-c4.json | ed3771e0a6f1b7933c870e61be62a89dc9501f9f820dd83f705b0bb15b4120cb |
| result-c16.json | 99fba85e51c6ce57ca05399aa2fbbc1802bd62939af0303b1d79be239e0b9ee7 |
| result-c32.json | c9dcdda0e8d751159af853b1692c19b9bc34f740fbf58535a44e3eea2b08fa09 |
| 10-table.md | ac07e74eacda7cbae3206f25634ff2b29931ddab9d662c711997128b35fafeb9 |
The L0 directory also has none of the run artefacts every other layer has (00-target.log, 03-apply-job.log, 04-nvidia-smi-load.csv.log, 08-job-logs.log), and its files carry a local modification time of 2026-09-28T09:06:22Z, about 7.5 minutes after the Stage C job finished writing results at 08:58:55Z. The brief asked for L0 to be re-confirmed on the exact cluster state used for L1 onward; that re-run did not happen. The single L0 measurement was taken 2026-09-28T08:46:33Z to 08:58:55Z on the same two nodes and the same head and worker pods that L1 later used (L1 ran on those pods after one restart caused by applying the policy; L3 ran on fresh pods on the same nodes).
Consequence: every L0 to L1 and L0 to L3 delta in Results: hardening compares runs from different sessions (different pod lifetimes, and for L3 different pod instances). With n=1 per level, the comparison cannot separate a layer effect from session-to-session variance.
Run timeline (UTC, 2026-09-28)
| Window | Activity | Evidence |
|---|---|---|
| 04:22 to 04:38 | Stage A inventory, neighbour scan, scanners (CPU only) | stageA-06-neighbor, stageA-08-scanners |
| 06:04:35 to 06:27:07 | Stage B benchmark (L4, 7B) | stageB-07-bench |
| 06:27 | Stage B neighbour re-probe | stageB-08-neighbor-reprobe |
| 08:46:33 to 08:58:55 | Stage C distributed baseline (the only L0 run) | stageC-06-bench |
| about 09:00 to 09:02 | Stage C cross-namespace re-inventory, Job API attempt, plaintext check | stageC-07-reinventory |
| 09:21:46 to 09:33:46 | L1 benchmark; L1 re-probe log written 09:33:48 | stageD-01-L1 |
| 09:43:18 to 09:44:19 | L2 Linkerd benchmark attempt (all requests failed) | stageD-02-L2 |
| 10:00:58 to 10:13:40 | L3 benchmark; auth proof log written 10:13:41 | stageD-03-L3 |
| about 10:27 | GPU on-demand group scaled 2 to 0 | stageD-99-gpu-teardown |
| 12:01 | GPU on-demand group scaled 0 to 2; 2 x g5.xlarge returned | stageD-04-L2-wireguard |
| 12:13:11 to 12:21:07 | L2b session: L0-before | bench-L0-before |
| 12:27:23 to 12:35:42 | L2b session: Cilium chaining plus WireGuard | bench-L2b |
| about 12:41 to 13:09 | Stale 05-cilium.conflist breaks new pod sandboxes | 22-stale-cni-conflist-breakage |
| 13:08:41 | Cilium reinstalled with encryption disabled and cni.uninstall=true | 23-cilium-reinstall-noenc |
| 13:09:03 | Stale conflist renamed on every node (same study team, parallel control session) | 22-remove-stale-cilium-conflist |
| 13:13:26 | Bracket-state check: Cilium agent running, Encryption: Disabled, no cilium_wg0, head pod not a Cilium endpoint | 24-bracket-state |
| 13:13:40 to 13:22:00 | L2b session: L0-after (Cilium agent still installed for the whole window) | bench-L0-after |
| 13:22:22 | Final Cilium uninstall, cni-uninstall=true | 30-cilium-uninstall-final |
| 13:24:54 | Zero GPU instances | 25-gpu-zero-verify |
Timestamps without seconds come from local file modification times converted to UTC, not from in-log # UTC stamps.
Safe-probing rules
- Probes targeted only pod IPs and Service names inside
sorami-lab-*namespaces, taken from an explicit allowlist built from the Kubernetes endpoints list (allowlist.txt). No CIDR sweeps. - Stage A used read-only handshakes only: HTTP GETs, TCP connects, and TLS client hellos. No POST to the Ray Job API.
- Stage C allowed exactly one synthetic Ray Job submission whose entrypoint printed a marker string and the hostname. No file writes, shell escapes, outbound network, or credential access.
- vLLM state-changing routes (
/scale_elastic_ep, response cancel, LoRA load/unload, sleep/wake) were enumerated from/openapi.jsonand not called. - KV and tensor channels: reachability and protocol identification only (connect, banner, TLS check). No injection, tampering, or decoding of model data.
Results: attack surface
Four default scanner configurations and the runtime surface
We rendered the vendor-default charts (KubeRay operator 1.7.1, ray-cluster 1.7.1, vLLM production-stack 0.1.12) into one 934-line manifest with 21 objects, including 1 RayCluster custom resource and 2 Deployments, and ran four scanners on it as in-cluster Jobs with their default settings. This section is about those four default configurations only. We did not write or test custom rules, admission-time policy engines that render custom resources, or other scanners. The properties in question (reachability from another namespace, raylet ports chosen at runtime, cleartext gRPC) are runtime and network properties, not necessarily violations of a static manifest rule, so a scanner that did not report them was not necessarily failing at its documented job.
| Scanner | Findings | Main themes | RayCluster mentions | Evidence |
|---|---|---|---|---|
| Trivy 0.67.0 | 31 of 143 tests failed (3 critical, 7 high, 10 medium, 11 low) | securityContext on router and operator, image tags and registry, operator ClusterRole on secrets, roles, services, networkpolicies | 0 | out-trivy |
| Checkov | 20 of 247 checks failed | securityContext, image digest and tag, SA token automount, CKV2_K8S_6 "pods which lack an associated NetworkPolicy" on the operator and router only | 0 | out-checkov |
| Kubescape v3.0.40 | 1 control (C-0013 non-root, 2 resources) | non-root containers; router and operator listed as highest-stake workloads | 0 | out-kubescape |
| kube-linter 0.8.3 | 5 lint errors | CPU request, dangling Service, :latest, read-only root FS, runAsNonRoot | 0 | out-kubelinter |
Fairness of comparing the finding counts with the RayCluster mentions (offline re-analysis of the stored scanner outputs and the rendered manifest, no new scans). What each scanner parsed, from its own output: Trivy's findings name the two Deployments, the KubeRay operator ClusterRole and two Roles; Checkov's 20 failed checks name the operator and router Deployments and the Pod objects it derives from them; Kubescape's one control covers 2 resources, and the two workloads it lists as highest-stake are the same two Deployments; kube-linter's 5 errors name 4 Deployment findings and one Service. Every finding therefore came from built-in workload, RBAC and Service kinds. None of the four outputs names the RayCluster custom resource, so we conclude the tools did not evaluate its embedded pod templates, not that they evaluated them and found nothing. Would any rule plausibly have applied to a RayCluster? The rendered RayCluster (69 lines in ray-cluster-1.7.1.yaml) sets no securityContext, references images by tag and not digest (rayproject/ray:2.52.0), sets imagePullPolicy: IfNotPresent and does not set automountServiceAccountToken. Rules of the kind that fired on the two Deployments (securityContext, non-root, read-only root filesystem, image digest, pull policy, service account token mounting, and Checkov's missing-NetworkPolicy graph check) would plausibly have fired on those pod templates if a tool had expanded them. That is an inference from rule wording and the manifest, not a measurement. No rule in the four outputs concerns reachability from another namespace, runtime listening ports, or cleartext gRPC, and the RayCluster declares no ports at all, so we would not expect a manifest rule to report the exposure in this section. The comparison is therefore not like for like: the finding counts in the table are hygiene findings on the objects the tools parsed, and zero counts mentions of an object they did not parse. It supports "the RayCluster was not evaluated", and it does not support a general claim about what scanners can detect. Whether a scanner with custom rules, or one that renders custom resources, would report the exposure was not tested.
The zero-mention result is a grep -aicE 'raycluster|ray-head|ray-worker|headGroupSpec|workerGroupSpecs' over each scanner output (crd-blindspot); we extended the same grep to the kube-linter output and it also returned 0. The scanners evaluated the two Deployments and the RBAC objects. The RayCluster, which is what creates the head and worker pods and the Ray listeners, was not evaluated. Checkov's missing-NetworkPolicy check fired for the operator and router but not for the Ray pods, because those pods exist only after the operator reconciles the custom resource.
A second, separate observation is visibility drift between declared port metadata and runtime listeners. The Stage A pod specs declare only containerPort 8080 (metrics) on the Ray head and worker (container-ports). A full TCP connect scan from the neighbour found far more (nmap-full-tcp): 15 of 17 observed Ray listening sockets were absent from the pod's declared port metadata. containerPort is metadata, not a security boundary, and Kubernetes does not block undeclared ports, so this is not a bypass of any Kubernetes control. The finding is narrower: a review or tooling pipeline that treats declared ports as an inventory of runtime exposure has incomplete visibility, seeing here a small fraction of the listeners that were reachable.
| Pod | Declared ports | Open ports seen by neighbour | Absent from declared metadata |
|---|---|---|---|
| Ray head | 8080 | 6379, 8080, 8265, 10001, 36087, 36349, 44217, 44227, 52365, 57914, 62226 (11) | 10 |
| Ray worker | 8080 | 8080, 36027, 36999, 45109, 52365, 57992 (6) | 5 |
| KubeRay operator | 8080 | 8080, 8082 | 1 |
| vLLM router | 8000, 9000 | 8000 | 0 |
A review that reasons from declared ports (manual review or port-based policy generators) sees 1 of 11 head ports. Stage C showed the same pattern on the GPU pipeline: the head listened on 23 TCP sockets and the worker on 15, while the manifest declared 4 head ports and none on the worker (02-head-listen-sockets, stageC-ray-pp.yaml).
No chart shipped any NetworkPolicy (netpol), all four Stage A pods automounted a ServiceAccount token, and none set runAsNonRoot (sa-automount). The KubeRay operator ClusterRole grants create, delete, get, list, update, and watch on secrets cluster-wide (kuberay-clusterrole-secrets); Trivy did flag this one (AVD-KSV-0041, critical).

Figure 1: The default scanners did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis. Bars show what each tool reported on the objects it did parse.
Exposed unauthenticated endpoints
All results below are from the unprivileged neighbour pod.
| Stage | Surface | Result | Auth | TLS | Evidence |
|---|---|---|---|---|---|
| A (CPU Ray) | Ray dashboard 8265 /api/version | 200, Ray 2.52.0, commit, session name | none | no | handshakes |
| A | Ray 8265 /api/jobs/ | 200, job list (empty) | none | no | same |
| A | Ray 8265 /nodes?view=summary | 200, 17,097 bytes: hostnames, IPs, CPU and memory per node | none | no | same |
| A | Ray GCS 6379, client 10001, ephemeral RPC | TCP accepted, non-TLS reply | not tested | no | same |
| A | Ray metrics 8080 head / worker | 200, 143,182 / 42,392 bytes | none | no | same |
| A | vLLM router /health, /v1/models, /metrics | 200 | none | no | stageA-baseline |
| B (L4 engine) | /v1/models, /v1/chat/completions, /v1/completions | 200, model listed, completions returned | none | no | 05-models, 06-chat, 07-completions |
| B | /openapi.json | 26 routes on port 8000, incl. /tokenize, /detokenize, /v1/embeddings, /v1/responses/{id}, /scale_elastic_ep (listed, not each called) | none configured | no | 13b-openapi-routes-parsed |
| B | /tokenize, /detokenize | text to 16 token ids; ids [785,6722,315,9625,374] back to "The capital of France is" | none | no | 08-tokenize, 09-detokenize |
| B | /metrics | 66 metric families incl. prefix_cache_hits_total 41,552 of prefix_cache_queries_total 46,357 | none | no | 11-metrics-counters, 12-metrics-names |
| C (PP=2) | vLLM GET /v1/models, POST /v1/completions | 200, 200 | none | no | 04-vllm-unauth |
| C | Ray GCS 6379, raylet 10002 to 10006 | TCP accepted, cleartext gRPC preface | none observed | no | 01-portscan, 03-plaintext |
| C | Ray Job API 8265 | connection refused | n/a | n/a | 02-rayjob-8265 |
The tokenizer routes turn any leaked token-id stream (logs, traces, or a captured tensor channel) back into text without credentials. The prefix-cache counters, combined with the ability to send prompts, suggest a cross-tenant probe of whether a guessed prefix was recently served. That side channel is inferred and was not tested.
The distributed control plane: cleartext and reachable
Wire evidence. From the neighbour, a raw TCP connect to Ray GCS 6379 on the head and to raylet 10002 on both head and worker returned 46 bytes immediately, beginning 00 00 18 04 00 00 00 00 00: an HTTP/2 frame header of length 0x18 (24 bytes), type 0x04 (SETTINGS), on stream 0. That is the server side of a cleartext HTTP/2 (h2c) connection, which is how gRPC runs without TLS. A TLS ClientHello to the same ports failed with SSL: WRONG_VERSION_NUMBER (03-plaintext, 01-plaintext-connect). A Redis PING to 6379 got the same SETTINGS frame rather than +PONG, confirming that modern Ray GCS is gRPC and not Redis (02-plaintext-gcs-ping).
Log excerpt: stageC-07-reinventory/03-plaintext.log (from the neighbour pod, byte samples shortened)
== plaintext peek ==
<head>:6379 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
<head>:10002 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
<worker>:10002 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
== TLS handshake attempt (expect refused) ==
<head>:6379 TLS_REFUSED SSLError [SSL: WRONG_VERSION_NUMBER] wrong version number
<head>:10002 TLS_REFUSED SSLError [SSL: WRONG_VERSION_NUMBER] wrong version numberObservation
Both Ray listeners opened with a cleartext HTTP/2 SETTINGS frame and could not negotiate TLS. No gRPC message was decoded.
We did not capture inter-node pipeline traffic with tcpdump, and we did not decode any gRPC messages. The claim is therefore narrower than "activations are readable on the wire": the Ray control and coordination listeners accept unauthenticated cleartext gRPC connections from the neighbour pod, and a TLS client cannot even negotiate with them. The NCCL activation channel between PP stages was not separately identified or captured, so its encryption state is inferred from NCCL's documented defaults (TCP sockets, no TLS) and was not measured.
Job API and remote code execution: claim by claim.
| Claim | Demonstrated? | Basis |
|---|---|---|
| The Ray Job API can be exposed to another namespace | Yes | Stage A: 8265 answered unauthenticated GETs from the neighbour pod (handshakes) |
| The default CPU Ray deployment exposed it | Yes | Stage A, KubeRay default image rayproject/ray:2.52.0: /api/jobs/ returned 200 with an empty list. No job was submitted there |
| The GPU vLLM deployment exposed it | No | Stage C: 8265 refused from the neighbour and from inside the head pod (02-rayjob-8265, 03-dashboard-http-inside) |
| Remote code execution against the GPU inference stack | No | The single permitted synthetic submission failed with connection refused. No code ran on any Ray node |
| Ray control ports were reachable cross-namespace | Yes | Stage C: 6379 and 10002 to 10006 open on the head, 10002 on the worker, from the neighbour pod (01-portscan) |
Job API result (negative for the GPU stack, image-specific). On the Stage A CPU RayCluster (KubeRay default image, ray[default]), 8265 answered unauthenticated GETs including /api/jobs/, so submission is very likely possible; per the safe-probing rules we did not POST there. On the Stage C GPU pipeline we made the one permitted synthetic submission: TCP to 8265 was refused and the submission failed with URLError [Errno 111] Connection refused (02-rayjob-8265). From inside the head pod, 127.0.0.1:8265 was also refused (03-dashboard-http-inside). The vllm/vllm-openai:v0.11.0 image ships minimal Ray, so the dashboard and job server never start. A cross-namespace Ray Client attempt on 10001 also failed (ray[client] absent; connection timeout) (02-crossns-ray-client-marker, 03-crossns-rayclient-10001).
Log excerpt: stageC-07-reinventory/02-rayjob-8265.log (the one permitted synthetic submission)
TCP <head>:8265 REFUSED (ConnectionRefusedError)
SUBMIT FAILED URLError <urlopen error [Errno 111] Connection refused>So remote code execution via the Job API was not demonstrated on the GPU pipeline, and was not attempted on the CPU deployment. Whether code execution is possible depends on the image: any image with ray[default], which is what the KubeRay default image used in Stage A ships, exposes 8265. Whether a client with network access to GCS and raylet alone can schedule work was not tested.

Figure 2: Declared versus observed listening ports. 10 of 11 head sockets and 5 of 6 worker sockets were undeclared.
Reachability matrix
Cross-namespace from the neighbour pod. "Open" means TCP connect succeeded.
Log excerpt: stageC-07-reinventory/01-portscan.log (head pod, L0)
=== target <head> ===
6379 :: OPEN
8000 :: OPEN
8265 :: closed
10001 :: closed
10002 :: OPEN
10003 :: OPEN
10004 :: OPEN
10005 :: OPEN
10006 :: OPEN| Target (Stage C/D) | Port | Purpose | L0 (Stage C) | After L1 | After L3 |
|---|---|---|---|---|---|
| Head (A10G) | 6379 | Ray GCS | open | blocked | not re-probed (L1 still in place) |
| Head | 8000 | vLLM API | open, no auth | open, no auth | open; ordinary requests to /v1/models and /v1/completions get 401 without key; other routes and the CVE-2026-48746 bypass not tested |
| Head | 10002 to 10006 | raylet, object manager | open | blocked | not re-probed |
| Head | 8265 | Ray Job API | closed | not probed (denied by policy) | not re-probed |
| Head | 10001 | Ray Client | closed | not probed | not re-probed |
| Worker (T4) | 10002 | raylet | open | blocked | not re-probed |
| Worker | 6379, 8000, 10003 to 10006 | not listening | closed | blocked | not re-probed |
Sources: 01-portscan, 01-reprobe-after-L1, 01-auth-proof. The L1 re-probe log lists 6379, 8000, and 10002 to 10006 on both hosts; it does not include 8265, which was not listening at L0 anyway. Note that the L1 probe reports "BLOCKED" on the worker for ports that were already closed at L0; the probe cannot distinguish a dropped SYN from a port that is not listening, so only the ports that were open at L0 (6379 and 10002 to 10006 on the head, 10002 on the worker) are evidence that L1 works. All probes were steady-state probes of long-running pods; the startup window after a new pod is created (Kubernetes networking and NetworkPolicy) was not probed.

Figure 3: Cross-namespace reachability of the Ray pods before and after the ingress NetworkPolicy.
Results: hardening
Single-node reference (Stage B, not comparable to Stage D)
For context only: one g6.xlarge (L4 24 GB), Qwen2.5-7B-Instruct bf16, vendor defaults, run 2026-09-28T06:04:35Z to 06:27:07Z. 128 of 128 requests succeeded at every level.
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.136 | 17.4 | 69.2 / 71.8 / 78.8 | 57.3 / 57.3 / 57.4 | 7344.9 / 7352.3 / 7354.6 |
| 4 | 0.531 | 67.9 | 121.9 / 122.9 / 153.3 | 58.4 / 58.5 / 58.5 | 7538.4 / 7546.8 / 7567.2 |
| 16 | 2.07 | 264.9 | 135.3 / 152.4 / 175.5 | 59.8 / 59.9 / 59.9 | 7725.3 / 7744.8 / 7765.4 |
| 32 | 3.807 | 487.2 | 173.9 / 192.6 / 217.5 | 64.7 / 64.9 / 65.1 | 8393.2 / 8430.3 / 8451.5 |
Source: stageB-07-bench result-c*.json. The L4 ran at 99% median utilisation and a median 71.9 W against its 72 W cap across 63 load samples (04-nvidia-smi-load). Tail spread was tight (p99 within 1% of p50 for E2E at every level), which contrasts with the two-node pipeline below.
Distributed per-layer results
Same two nodes, same model and flags, same client. 128 requests per level. L0, L1, and L3 each had 0 failures across 512 requests. All numbers are from the per-level result-c*.json files.
L0: distributed baseline, no hardening (stageC-06-bench, reused as stageD-00-L0).
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.269 | 34.4 | 38.2 / 47.7 / 51.4 | 28.9 / 31.3 / 32.7 | 3712.4 / 4026.4 / 4186.3 |
| 4 | 0.927 | 118.7 | 59.7 / 386.4 / 7397.3 | 32.1 / 34.1 / 34.3 | 4138.9 / 4497.6 / 11748.2 |
| 16 | 3.198 | 409.4 | 81.5 / 798.5 / 800.8 | 37.3 / 41.7 / 41.7 | 4807.4 / 5795.4 / 5797.1 |
| 32 | 3.353 | 429.2 | 474.6 / 10024.8 / 10028.8 | 54.0 / 56.7 / 131.1 | 7330.0 / 16881.9 / 16899.2 |
L1: NetworkPolicy (stageD-01-L1).
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.271 | 34.6 | 38.0 / 47.9 / 49.4 | 28.7 / 31.0 / 31.8 | 3685.8 / 3979.5 / 4074.5 |
| 4 | 1.011 | 129.5 | 60.4 / 91.9 / 405.4 | 30.7 / 32.4 / 32.8 | 3954.7 / 4176.5 / 4572.8 |
| 16 | 3.109 | 398.0 | 92.7 / 986.9 / 988.2 | 37.2 / 42.5 / 42.5 | 4801.9 / 6079.6 / 6082.1 |
| 32 | 4.444 | 568.8 | 165.1 / 837.0 / 839.3 | 54.1 / 54.5 / 57.4 | 7031.6 / 7760.3 / 7776.2 |
L3: NetworkPolicy plus engine API key (stageD-03-L3).
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.268 | 34.3 | 38.6 / 48.4 / 51.5 | 29.0 / 31.5 / 32.3 | 3717.9 / 4044.7 / 4138.7 |
| 4 | 0.952 | 121.9 | 61.1 / 385.8 / 7405.5 | 30.8 / 32.5 / 32.7 | 3970.0 / 4487.8 / 11537.1 |
| 16 | 2.555 | 327.0 | 73.4 / 10606.2 / 10624.3 | 37.0 / 40.3 / 41.8 | 4776.7 / 15584.5 / 15604.3 |
| 32 | 4.346 | 556.2 | 218.0 / 786.5 / 1028.2 | 54.0 / 61.5 / 61.5 | 7174.6 / 8597.5 / 8599.4 |
GPU utilisation under load (samples every ~20 s; min / median / max):
| Layer | Samples | Head A10G util % | Worker T4 util % | Head / worker peak memory MiB |
|---|---|---|---|---|
| L0 | 34 | 0 / 12 / 15 | 0 / 37 / 68 | 14,160 / 8,583 |
| L1 | 33 | 0 / 12 / 16 | 0 / 39 / 59 | 14,160 / 8,557 |
| L3 | 35 | 0 / 12 / 16 | 0 / 37 / 67 | 14,140 / 8,541 |
Sources: 04-nvidia-smi-load.csv.log in each layer directory. The T4 (PP stage 1) is consistently the busier device and is the likely pipeline bottleneck. Neither GPU approaches saturation, which fits a small model in eager mode where per-step scheduling and inter-stage transfer, not compute, dominate. GPU utilisation is indistinguishable across layers.
Performance comparison
Percent change versus L0, computed from the raw JSON. Positive on throughput means higher than L0; positive on latency means slower. These are single runs from different sessions (Baseline provenance (important)). No confidence intervals can be computed from n=1, and the numbers must be read as "observed difference", not "effect of the layer".
| conc | Layer | out tok/s | TTFT p50 | TPOT p50 | E2E p50 | E2E p95 |
|---|---|---|---|---|---|---|
| 1 | L1 | +0.6% | -0.5% | -0.7% | -0.7% | -1.2% |
| 1 | L3 | -0.3% | +1.0% | +0.3% | +0.1% | +0.5% |
| 4 | L1 | +9.1% | +1.2% | -4.4% | -4.5% | -7.1% |
| 4 | L3 | +2.7% | +2.3% | -4.0% | -4.1% | -0.2% |
| 16 | L1 | -2.8% | +13.7% | -0.3% | -0.1% | +4.9% |
| 16 | L3 | -20.1% | -9.9% | -0.8% | -0.6% | +168.9% |
| 32 | L1 | +32.5% | -65.2% | +0.2% | -4.1% | -54.0% |
| 32 | L3 | +29.6% | -54.1% | 0.0% | -2.1% | -49.1% |
| any | L2 (Linkerd) | failed: 0 of 128 ok at every level | n/a | n/a | n/a | n/a |
| 1 | L2b (Cilium chaining + WireGuard), vs mean of same-session L0 | -3.5% | +8.3% | +3.6% | +3.2% | +3.1% |
| 4 | L2b | -3.4% | +6.2% | +3.5% | +3.8% | +4.0% |
| 16 | L2b | -4.5% | +11.1% | +5.3% | +5.4% | +3.4% |
| 32 | L2b | -6.5% | +45.8% | +8.4% | +8.4% | +4.3% |
The L2b rows use a different baseline and different hardware (L2b: Cilium chaining plus WireGuard combined (single bracketed run)): they compare against the mean of two same-session L0 runs on 2 x A10G and are not comparable in absolute terms to the L1 and L3 rows. The L2b TTFT p50 deltas are small absolute changes (for example 66.7 to 97.2 ms at c=32) and the L0-to-L0 TTFT drift at c=32 was -24.1%, so they are not read as a difference.
How to read this table (L1 and L3 rows).
- Median per-token and median full-request (E2E) latency are stable. TPOT p50 is within 4.4% of L0 at every level for both layers, and E2E p50 is within 4.5%. This is the most stable signal in the data, because medians over 128 samples are insensitive to a handful of stalled requests.
- Throughput and tail latency swing in both directions by large amounts. At c=32 both hardened layers show about 30% more throughput than L0; at c=16 L3 shows 20% less. Neither NetworkPolicy (per-packet eBPF lookup) nor a bearer-token string comparison is a plausible cause of a 30% gain or a 20% loss when TPOT p50 moved by under 1%.
- The swings trace to a small number of stalled requests. L0 at c=32 has TTFT p95 of 10,024.8 ms against a p50 of 474.6 ms, and L3 at c=16 has TTFT p95 of 10,606.2 ms against a p50 of 73.4 ms; the same pattern of multi-second TTFT outliers appears at c=4 in L0 (p99 7,397.3 ms) and L3 (p99 7,405.5 ms) but not in L1. Because the harness is closed-loop and the wall time at c=16 and c=32 is only 29 to 50 s, a few requests that wait roughly 7 to 10 s before their first token stretch the wall clock and drag throughput down for the whole level. We did not identify the cause of these stalls (candidates include scheduler admission under PP, Ray actor scheduling, and eager-mode first-step cost); the stall pattern is independent of which layer is applied.
- The L0 run at c=32 is the outlier in this study. It happens to be the reused baseline, so every c=32 comparison inherits its stall.
Conclusion for RQ3 (performance). Within cross-session variance, with n=1 run per level, we observed no measurable steady-state difference for L1 (NetworkPolicy) or L3 (NetworkPolicy plus API key). We do not claim that either layer improves throughput. For L2b (Cilium chaining plus WireGuard combined), measured against a same-session bracket, the observed difference in this single bracketed run was 3.4% to 6.5% lower throughput and 3.2% to 8.4% higher median E2E; the c=1 and c=4 part is within L0-to-L0 drift, and the c=16 and c=32 throughput difference is outside it. With n=1 per condition, no per-request data, no confidence intervals, heterogeneous GPUs across sessions and a closed-loop client, we do not present these figures as a measured cost of encryption. The data cannot rule out a difference smaller than the run-to-run spread, which on this topology is at least several percent on median E2E and tens of percent on throughput and tail latency.
Security effect of each layer
L1: NetworkPolicy. After applying the three policies, the same probe script from the same neighbour pod turned every port that was open at L0 into BLOCKED, while 8000 stayed open as intended (01-reprobe-after-L1).
Command: apply the L1 policies (Stage D)
kubectl apply -f manifests/stageD-L1-networkpolicy.yamlLog excerpt: stageD-01-L1/01-reprobe-after-L1.log (head pod, same probe from the same neighbour pod)
6379 :: BLOCKED
8000 :: OPEN
10002 :: BLOCKED
10003 :: BLOCKED
10004 :: BLOCKED
10005 :: BLOCKED
10006 :: BLOCKED| Target | Port(s) | L0 | L1 |
|---|---|---|---|
| Head | 6379 | open | blocked |
| Head | 10002 to 10006 | open | blocked |
| Head | 8000 | open | open (allowed) |
| Worker | 10002 | open | blocked |
Operational effect: applying default-deny to the running Ray cluster forced one restart of both engine pods; the L1 target log shows RESTARTS 1 (3m41s ago) on the head and 1 (3m42s ago) on the worker at 09:21:46Z (00-target). The cluster re-formed under the policy and the benchmark ran clean. The mechanism of the restart (for example conntrack state for established Ray connections being dropped when the eBPF policy attached) was not investigated.
Observation
What L1 does not do: it does not authenticate callers on 8000, it does not encrypt anything, and it allows all traffic between engine pods, so a compromise of either engine pod still reaches the whole Ray control plane. It is ingress-only (policyTypes: ["Ingress"], line 22 of manifests/stageD-L1-networkpolicy.yaml; egress left open by design per the comment at line 11), so it does not show containment of a compromised engine: outbound connections from the engine to other namespaces, the instance metadata service, or the internet were neither restricted nor tested. The evidence is steady-state only; the VPC CNI enforcing mode was not recorded (Limitations), so a possible default-allow window for newly started pods is unmeasured.
L3: API key on top of L1. The engine was restarted with --api-key sourced from a Kubernetes Secret (01-auth-proof):
Log excerpt: stageD-03-L3/01-auth-proof.log
--- no auth ---
no-key /v1/models -> 401
no-key /v1/completions -> 401
--- wrong key ---
bad-key /v1/models -> 401
--- correct key ---
good-key /v1/completions -> 000
command terminated with exit code 28
good-key /v1/models -> 200
good-key /v1/completions -> 200| Call from neighbour | Before (Stage C) | After L3 |
|---|---|---|
No key, GET /v1/models | 200 | 401 |
No key, POST /v1/completions | 200 | 401 |
Wrong key, GET /v1/models | 200 | 401 |
Correct key, GET /v1/models | n/a | 200 |
Correct key, POST /v1/completions | n/a | 200 (a first attempt timed out with curl exit 28, the retry returned 200) |
Observation
What this shows, and what it does not. L3 is the weakest security result in this study. It shows that the normal key path returned 401 on two /v1 routes (/v1/models, /v1/completions) with no key or a wrong key, and 200 with the key. It does not show that API-key authentication protects vLLM, and it is not evidence of authentication closure, for three reasons:
- Route coverage. We did not test the other 24 routes. In vLLM the check applies to the
/v1routes, and/metrics,/health,/tokenize, and/detokenizemay remain open. That is unverified here. - Known bypass. The tested engine was
vllm/vllm-openai:v0.11.0, which is inside the affected range (>=0.3.0, <0.22.0) of CVE-2026-48746 / GHSA-94f4-hr76-p5j6 (CVSS 9.1, published 2026-06-02): a craftedHostheader makesAuthenticationMiddlewarereconstruct a path that skips the/v1API-key check [19] [C]. The advisory says deployments behind an RFC-conforming proxy such as nginx are not affected; ours was exposed directly on its Service with no proxy. The advisory was public on the experiment date (2026-09-28), so the tested version was publicly known to be bypassable at experiment time. We did not test the bypass. - Cleartext key. The key travels in cleartext over plain HTTP, so any party that can observe pod traffic can capture it.
Any reuse of this configuration should run vLLM 0.22.0 or later, or put an RFC-conforming proxy in front, and then test every route with no key, a wrong key, and a correct key.
L2: encryption in transit
L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)
We treat this section as an engineering incident report, not a security finding: it says nothing about the attack surface in Results: attack surface. The intent was to encrypt and mutually authenticate the client to engine path with a service mesh, leaving everything else unchanged. Linkerd (edge channel; exact version not kept, and proxy logs not kept) installed cleanly: the destination, identity, and proxy-injector pods were Running, and both engine pods were meshed with an issued workload identity default.sorami-lab-stack-vllm.serviceaccount.identity.linkerd.cluster.local. Ray ports 6379 and 10002 to 10010 were excluded from the proxy because routing them through it deadlocked raylet on this build. With those exclusions ray status showed both nodes active, and a completion issued from inside the head pod worked (05-l2-outcome).
The benchmark client was then meshed too (linkerd.io/inject: enabled plus proxy-await) so that the client to 8000 hop would be mTLS end to end. Every request failed:
| conc | ok / requests | wall s | error |
|---|---|---|---|
| 1 | 0 / 128 | 0.36 | HTTPError 503: Service Unavailable |
| 4 | 0 / 128 | 0.13 | same |
| 16 | 0 / 128 | 0.13 | same |
| 32 | 0 / 128 | 0.13 | same |
Source: stageD-02-L2 result-c*.json, 08-job-logs. The 0.13 s wall time for 128 requests means the client-side proxy rejected requests immediately; they never reached the engine. The operator notes record that the L7 HTTP route returned "route default.http: service unavailable", and that switching port 8000 to opaque (L4) mode produced connection resets before the inbound proxy on the head saw any traffic (05-l2-outcome). Those two observations come from the outcome note rather than from captured proxy logs, and the opaque-mode run did not produce a result file of its own.
Interpretation, with the limits stated. The tested Linkerd configuration failed to carry the vLLM workload; the root cause was not isolated. This is not evidence that Linkerd is incompatible with vLLM or with SSE in general: the Linkerd version and the proxy logs were not kept, Service versus pod-IP routing was not separated, and one-sided versus two-sided meshing was not compared, so the failure cannot be attributed to SSE, to vLLM, or to Linkerd as a technology. Plausible contributors include the interaction between the Service used by the client (the KubeRay-managed -openai Service) and Linkerd's endpoint discovery for pods whose Ray ports were skipped, protocol detection on port 8000, and the ordering of sidecar readiness against vLLM readiness. The finding we are confident in is narrow and operational: a team that plans to mesh a Ray-backed vLLM deployment should expect to debug it, and should not assume that "inject the proxy" preserves the serving path. Because the meshed path never carried a successful request, no mTLS performance difference is reported. We also do not report a literature figure in its place.
Security effect of L2a as deployed: none on the attack surface measured in Results: attack surface. The Ray control-plane ports were excluded from the proxy, so even a working L2a would have left GCS and raylet traffic in cleartext.
L2b: Cilium chaining plus WireGuard combined (single bracketed run)
Node-level WireGuard is designed to encrypt pod traffic that crosses nodes without touching the application or its ports. Here it was tested only as part of the combined Cilium chaining plus WireGuard configuration (L2b). Interface counters showed inter-node pod traffic carried by the WireGuard device. Ray GCS, raylet, object transfer and NCCL flows were not individually identified or captured, so that these flows are encrypted is an inference from the datapath, not a measurement. It was not installable in place during the main session, so it ran later the same day as a separate, scoped experiment (stageD-04-L2-wireguard, decision record 02-decision).
Setup. The VPC CNI addon v1.22.4-eksbuild.3 exposes no encryption setting (01-vpccni-encryption-options). We installed Cilium 1.20.2 in CNI chaining mode behind aws-cni (cni.chainingMode=aws-cni, native routing, no masquerade, kubeProxyReplacement=false, policyEnforcementMode=never, encryption.type=wireguard, nodeEncryption=false, enableRouteMTUForCNIChaining=true). VPC CNI kept IPAM and pod IPs; only the recreated engine pods became Cilium endpoints (10-cilium-install, 11-recreate-engine-under-cilium). The L1 policies were removed for all three runs and restored afterwards, and no API key was set, so encryption was the only intended variable.
Hardware change. On-demand capacity returned 2 x g5.xlarge (A10G + A10G) rather than the Stage C g5 + g4dn pair. Head and worker were pinned to the same two nodes (ip-10-99-17-131, ip-10-99-38-120) for all three runs. Removing the T4 bottleneck roughly tripled c=32 throughput (about 1180 against 429 tok/s), so absolute L2b numbers are not comparable to L0, L1, or L3 in Distributed per-layer results. Only the within-session deltas below are meaningful.
Order (UTC, 2026-09-28; one run per level, 128 requests per level, 128 of 128 ok at every level in all three runs):
| Run | Bench window | State |
|---|---|---|
| L0-before | 12:13:11 to 12:21:07 | plain VPC CNI, no encryption |
| L2b | 12:27:23 to 12:35:42 | Cilium chained + WireGuard |
| L0-after | 13:13:40 to 13:22:00 | Cilium agent reinstalled and running with encryption disabled, no cilium_wg0; head pod not a Cilium endpoint; worker pod not checked (24-bracket-state). Not a proven Cilium-free baseline (see below) |
L0-before (bench-L0-before):
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.397 | 50.8 | 29.2 / 33.6 / 38.7 | 19.6 / 20.4 / 20.6 | 2522.0 / 2619.5 / 2655.0 |
| 4 | 1.458 | 186.7 | 40.5 / 56.1 / 105.0 | 21.2 / 22.0 / 22.2 | 2730.1 / 2830.7 / 2860.3 |
| 16 | 5.313 | 680.1 | 51.8 / 88.2 / 100.4 | 23.1 / 23.9 / 24.0 | 2984.8 / 3102.8 / 3104.9 |
| 32 | 9.217 | 1179.8 | 75.8 / 144.2 / 149.5 | 26.6 / 27.9 / 27.9 | 3434.3 / 3615.4 / 3617.3 |
L2b, Cilium chaining plus WireGuard (bench-L2b):
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.381 | 48.7 | 31.4 / 35.0 / 37.4 | 20.4 / 21.2 / 21.7 | 2618.4 / 2720.4 / 2783.1 |
| 4 | 1.400 | 179.2 | 43.4 / 65.8 / 117.5 | 22.1 / 23.0 / 23.2 | 2853.9 / 2967.0 / 2991.4 |
| 16 | 4.986 | 638.2 | 60.6 / 80.3 / 100.5 | 24.9 / 25.2 / 25.3 | 3216.2 / 3264.9 / 3270.0 |
| 32 | 8.726 | 1116.9 | 97.2 / 125.8 / 158.1 | 28.4 / 28.8 / 28.9 | 3665.8 / 3759.8 / 3763.1 |
L0-after, drift bracket (bench-L0-after):
| conc | req/s | out tok/s | TTFT p50 / p95 / p99 ms | TPOT p50 / p95 / p99 ms | E2E p50 / p95 / p99 ms |
|---|---|---|---|---|---|
| 1 | 0.392 | 50.1 | 28.8 / 33.3 / 36.0 | 19.8 / 20.7 / 21.0 | 2550.2 / 2659.0 / 2696.5 |
| 4 | 1.441 | 184.4 | 41.2 / 49.6 / 116.4 | 21.5 / 22.3 / 22.4 | 2766.7 / 2872.7 / 2880.1 |
| 16 | 5.131 | 656.8 | 57.3 / 90.9 / 100.0 | 24.2 / 24.9 / 24.9 | 3119.8 / 3210.1 / 3214.2 |
| 32 | 9.445 | 1209.0 | 57.5 / 150.7 / 178.7 | 25.8 / 27.1 / 27.2 | 3326.8 / 3596.7 / 3626.6 |
Difference: L2b against the mean of the two L0 runs. Drift: L0-after against L0-before.
| conc | out tok/s L0 mean | L2b | difference | L0 drift | E2E p50 L0 mean ms | L2b ms | difference | L0 drift |
|---|---|---|---|---|---|---|---|---|
| 1 | 50.45 | 48.7 | -3.5% | -1.4% | 2536.1 | 2618.4 | +3.2% | +1.1% |
| 4 | 185.55 | 179.2 | -3.4% | -1.2% | 2748.4 | 2853.9 | +3.8% | +1.3% |
| 16 | 668.45 | 638.2 | -4.5% | -3.4% | 3052.3 | 3216.2 | +5.4% | +4.5% |
| 32 | 1194.4 | 1116.9 | -6.5% | +2.5% | 3380.6 | 3665.8 | +8.4% | -3.1% |
TPOT p50 difference on the same basis is +3.6%, +3.5%, +5.3%, and +8.4%. The full metric-by-metric table, including the difference against L0-before alone, is in 50-wireguard-tax.
How to read this. The L0-to-L0 drift reaches about 3.4% on throughput (at c=16), so the c=1 and c=4 difference (-3.5% and -3.4%) is within drift, and the c=16 and c=32 difference (-4.5% and -6.5%) is outside it. On E2E p50 the drift is larger at c=16 (+4.5%), so only the c=32 latency difference (+8.4%) is clearly outside drift. The direction is consistent: L2b is slower than both L0 runs at every level on throughput, E2E p50, and TPOT p50. Tail percentiles show no clean signal (TTFT p95 at c=32 is lower under L2b than in either L0 run), so we claim no tail-latency difference. GPU medians were unchanged across the three runs (head 18%, worker 22 to 23%; 04-nvidia-smi-load.csv.log in each directory), so neither GPU was saturated and an inter-node network effect can surface in TPOT.
Confound. The L2b run had Cilium's eBPF datapath chained behind aws-cni plus WireGuard, so the measured difference is Cilium chaining plus WireGuard combined, not WireGuard alone. Nothing here isolates WireGuard. We tried a Cilium-without-encryption state, but helm upgrade --set encryption.enabled=false did not restart the agents and cilium-dbg still reported WireGuard active, so we discarded that state and did not benchmark it (17-cilium-disable-wg, 18-recreate-engine-cilium-noenc). The Cilium-only share of the difference was not isolated.
Threat to the L2b comparison: the closing bracket is not a proven Cilium-free baseline. 24-bracket-state (13:13:26Z) shows the Cilium 1.20.2 agent reinstalled and running during L0-after, with CNI Chaining: aws-cni, Encryption: Disabled, and no cilium_wg0 interface. Cilium was not removed until 30-cilium-uninstall-final at 13:22:22Z, after the L0-after sweep ended at 13:22:00Z. The head pod IP was absent from the Cilium endpoint list, which is consistent with the stale conflist having been renamed at 13:09Z so new pods used 10-aws.conflist, but the worker pod was not checked. So L0-after is a no-encryption control whose datapath state is only partly verified. If the worker ran as a Cilium endpoint, L0-after contains part of the Cilium datapath effect, and averaging it into the baseline would shrink the apparent difference. The L2b difference against L0-before alone is in 50-wireguard-tax and is larger at c=16 (-6.2% throughput).
Encryption evidence. Agent status on both GPU nodes reported Encryption: Wireguard [NodeEncryption: Disabled, cilium_wg0 ... Port: 51871, Peers: 3] and CNI Chaining: aws-cni, with recent handshakes to every peer (12-cilium-encryption-status). We then sent 50 MiB of a repeated ASCII marker from the head pod to the worker pod (13-wg-counter-proof):
| Head-node counter | Before | After | Delta |
|---|---|---|---|
| Bytes sent by head pod, received by worker pod | n/a | 55,189,000 / 55,189,000 | 55.19 MB |
cilium_wg0 tx_bytes | 651,048 | 56,438,648 | +55.79 MB |
ens5 tx_bytes | 641,459,248 | 697,520,489 | +56.06 MB |
The payload went into the WireGuard device (about 0.6 MB above payload, consistent with TCP/IP headers). Over the L2b sweep the head cilium_wg0 tx grew 572.5 MB and the worker cilium_wg0 rx grew 546.7 MB, so the pipeline traffic also crossed the tunnel (14-wg-counters-before-bench, 15-wg-counters-after-bench). This is interface-counter evidence that traffic was routed through cilium_wg0. It is not a packet capture showing ciphertext on the ENI: the Cilium agent image has no tcpdump and no privileged node shell was available, so the absence of the cleartext marker or HTTP/2 SETTINGS preface on ens5 was not observed directly.
Security scope. WireGuard protects the node-to-node wire. It does not close the cross-namespace reachability from Reachability matrix, because a neighbour pod's connection to 6379 is decrypted before delivery. L2b complements L1 and does not replace it.
Operational finding: an incomplete Cilium removal procedure broke new pods. This is a gap in our removal procedure, not Cilium misbehaving. Cilium was installed with the chart default cni.uninstall=false, which is documented in the Cilium Helm reference as controlling whether the CNI configuration is removed on agent shutdown [20] [C]. helm uninstall cilium at 12:41:16Z removed the agents but left /etc/cni/net.d/05-cilium.conflist on all four nodes, including the two CPU nodes that never ran an engine pod (19-cilium-uninstall, 22-remove-stale-cilium-conflist). Because it sorts before 10-aws.conflist, the container runtime kept invoking cilium-cni, and every new pod sandbox failed with unable to connect to Cilium agent ... dial unix /var/run/cilium/cilium.sock: connect: no such file or directory (22-stale-cni-conflist-breakage). Existing pods kept running. The breakage lasted from about 12:41Z to 13:09Z, when the file was renamed on each node. A later install with cni.uninstall=true uninstalled cleanly at 13:22:22Z, and a sanity pod on a CPU node networked afterwards (30-cilium-uninstall-final, 31-cni-sanity-pod). The lesson is about security-control lifecycle testing: removal and rollback are part of the control. Anyone trialling Cilium chaining on EKS needs an explicit cleanup step: set cni.uninstall=true and confirm the agents restarted with it before uninstalling, or remove the conflist on every node afterwards and verify with a throwaway pod per node pool. Separately, flipping encryption.enabled with helm upgrade changed the ConfigMap but not the running agents; confirm state with cilium-dbg status after any toggle.
Teardown. The GPU on-demand group was scaled from 2 to 0 (23-tf-plan-gpu0, 24-tf-apply-gpu0); the instance count reached 0 at 13:24:54Z (25-gpu-zero-verify), and both GPU groups were confirmed at desired 0 at 13:28:54Z (41-gpu-zero-verify). The L0-after sweep finished at 13:22:00Z, before any GPU instance drained. The L1 policies were restored at 13:23:06Z (32-restore-L1-netpol). The conflist rename (13:09Z) and the GPU scale-down (13:22 to 13:24Z) were done by the same study team, in a parallel control session alongside the one that ran the benchmarks; both are recorded in 60-encryption-proof-and-outcome. Neither overlaps a benchmark window.

Figure 4: Successful requests per concurrency level for the two encryption attempts.

Figure 5: The Cilium chaining plus WireGuard throughput difference against the drift between the two same-session baseline runs.
Summary of layers
| Layer | Closes | Proof | Observed difference (n=1) | Leaves open |
|---|---|---|---|---|
| L1 NetworkPolicy (ingress only) | Cross-namespace ingress to Ray GCS and raylet from the tested neighbour pod (steady state) | 7 ports open at L0 (6 head, 1 worker) to blocked | no measurable steady-state difference within variance; one pod restart when applied live | Unauthenticated 8000; cleartext everywhere; lateral movement between engine pods; all egress (no containment of a compromised engine); pod startup window unmeasured |
| L2a Linkerd mTLS | Intended: encryption and identity on 8000 | identity issued, pods meshed | not measurable: the tested configuration failed to carry the workload (503), root cause not isolated | Everything, as deployed; Ray ports excluded |
| L2b Cilium chaining plus WireGuard combined | Cleartext inter-node pod traffic (Ray and NCCL flows inferred, not separately captured) | agent status, peers, and cilium_wg0 counters (+55.79 MB for a 55.19 MB transfer); no ciphertext pcap | 3.4% to 6.5% lower throughput, 3.2% to 8.4% higher median E2E (single bracketed run, n=1; c=1/c=4 within drift; closing bracket not proven Cilium-free) | Cross-namespace reachability (decrypted before delivery); callers not authenticated; removal needs an explicit CNI cleanup step |
| L3 API key (plus L1) | Ordinary unauthenticated requests on two /v1 routes (/v1/models, /v1/completions) | 200 to 401 without key, 200 with key | no measurable difference within variance | Not authentication closure: 24 other routes untested; v0.11.0 is in the CVE-2026-48746 bypass range (not tested); key sent in cleartext |
Discussion
Practical recommendations
Ordered by closed exposure per unit of effort, based on what this study measured.
- Apply namespace default-deny ingress before the engine starts. L1 was the only layer that blocked cross-namespace ingress to the Ray control plane from the tested neighbour, it showed no measurable steady-state difference, and it is plain Kubernetes. Allow the engine pods to reach each other on all ports (Ray picks ports at runtime) and allow the serving port only from named caller namespaces. Apply it with the RayCluster, not onto a live one, to avoid the restart seen here. Add an egress policy too if containment of a compromised engine matters; L1 here was ingress-only and gives no containment. On EKS, consider
NETWORK_POLICY_ENFORCING_MODE=strictso new pods start default-deny; the startup window in the defaultstandardmode was not measured here. - Turn on Ray's own controls where the version supports them. Ray 2.52.0 already had token authentication (
RAY_AUTH_MODE=token, off by default) covering the dashboard and GCS, and Ray supports TLS for gRPC. We did not evaluate either, so their effect on serving and their coverage are open questions, but they address the control-plane exposure at the application layer rather than relying on network placement alone. - Encrypt below the application. The multi-node surface is many ports, most absent from declared port metadata, speaking cleartext gRPC and NCCL. A sidecar mesh that skips those ports (as L2a had to) does not encrypt them. Node-level encryption (WireGuard in the CNI) is designed to cover them without per-port configuration; here the combined Cilium chaining plus WireGuard configuration routed inter-node pod traffic through the tunnel (Ray and NCCL flows were not separately captured) and carried SSE streaming where the tested L7 sidecar configuration failed, and in this single bracketed run the encrypted configuration showed about 3% to 8% lower throughput and higher median latency on this PP=2 topology (Cilium chaining plus WireGuard combined, n=1). On EKS that currently means chaining Cilium behind the VPC CNI, which needs an explicit CNI cleanup step on removal (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
- Choose the image deliberately. Whether the Ray Job API exists depends on whether the image ships
ray[default]. The minimal vLLM image did not; the KubeRay default image did. Treat an exposed unauthenticated Ray Job API on 8265 as a potential arbitrary-code-execution interface, and put it behind authentication or disable the dashboard. For the CPU deployment this rests on the documented Job API semantics; no job was submitted there. - Inventory from the running pod, not from the manifest. Declared port metadata covered 1 of 11 listening ports on the Stage A Ray head. Use a socket listing inside the pod or an allowlist-driven connect scan from a neighbour namespace.
- Require an API key on the engine, stored in a Secret, on a vLLM version that is not affected by CVE-2026-48746 (0.22.0 or later), or behind an RFC-conforming reverse proxy. In this study the normal key path returned 401 on two
/v1routes; that does not show the key protects vLLM, and the version tested was publicly known to be bypassable. Test every route (including/metrics,/tokenize,/detokenize) with no key, a wrong key, and the correct key before relying on it. - Scope the operator. The KubeRay operator ClusterRole with cluster-wide
secretsaccess is the most privileged identity in the stack. Namespace-scoped operator deployment, where supported, reduces that blast radius.
Defence in depth ordering
Each layer covers a different gap and none is sufficient alone.
| Question an attacker asks | Layer that answers it | Measured here |
|---|---|---|
| Can I reach it? | NetworkPolicy (L1) | yes, ingress from the tested neighbour, steady state |
| Can it reach out if compromised? | Egress NetworkPolicy | not tested; L1 was ingress-only |
| Will it serve me? | Authentication (L3), Ray token auth | not established: the normal key path returned 401 on two /v1 routes; other routes and the CVE-2026-48746 bypass untested; Ray token auth not evaluated |
| Can I read or alter it on the wire? | Encryption (L2) | the tested L2a configuration failed (engineering incident); the combined L2b configuration routed inter-node pod traffic through WireGuard (counter evidence, no pcap; NCCL not separately captured) |
| Can I run code through it? | Image choice, Job API auth, RBAC | Job API not listening on the GPU image, so no code execution demonstrated there (The distributed control plane: cleartext and reachable); not tested elsewhere |
L1 first, because it is cheap and it is the only tested layer that shrinks the reachable Ray surface without touching Ray. Encryption and Ray's own controls next, because L1 still allows the intended callers and anything that shares their namespace. An API key is a weaker, supporting layer: the normal key path returned 401, the version and route-coverage caveats above apply, and without encryption the key itself is sent in cleartext on the same network that L1 only partially fences.
Network-layer versus L7 encryption for streaming inference
The two L2 attempts split along the layer they operate at. The Linkerd sidecar sits in the HTTP path and has to understand, route, and hold open every SSE response; in the tested configuration it returned 503 for every request (root cause not isolated), and it had to skip the Ray ports to keep raylet alive. WireGuard in the CNI (tested only as the combined Cilium chaining plus WireGuard configuration) operates below TCP and does not parse the application protocol; it carried every streaming request, and interface counters showed inter-node pod traffic crossing the tunnel without per-port configuration (Ray and NCCL flows were not separately captured, so their encryption is inferred). The observed difference in this single bracketed run was about 3% to 8% on throughput and median latency, larger at higher concurrency where more activation bytes cross nodes per second. Three caveats bound this conclusion: the measured difference includes the Cilium chaining datapath, the closing L0 bracket was not proven Cilium-free, and with a 1.5B model the bytes moved between stages per decode step are small, so a larger model will likely pay more (inference, not measured).
What the four default scanner configurations did not report, and why
Three structural reasons are consistent with the outputs.
- Custom resources were not parsed. The four outputs name only built-in workload, RBAC and Service kinds. A RayCluster is a template for pods that the operator creates later, and none of the four tools mentioned it, so its pod spec and security context were not evaluated in these configurations.
- Runtime ports are not in the metadata. Ray and NCCL bind most listeners at runtime. A tool that reads only
containerPortcannot know they exist.containerPortis metadata, not a security boundary, and Kubernetes does not block undeclared ports. - Network posture is reported per object, not per flow. Checkov's CKV2_K8S_6 notes that a pod lacks a NetworkPolicy, which is useful, but no scanner output reported that a Ray control-plane listener was reachable from another namespace. That is a runtime property, and a reachability test such as the neighbour probe measures it.
These are observations about four tools at the versions tested, in default configurations, on one manifest. Custom rules, other scanners, newer versions, and admission-time policy engines that render custom resources may behave differently; we did not test them.
Limitations
Internal validity.
- Reused baseline. The Stage D L0 files are byte-identical copies of the Stage C benchmark (SHA-256 identical for all five files; no run logs in the L0 directory). L0 was not re-run in the same session as L1 or L3. All L1 and L3 deltas are cross-session. L2b is the exception: it has fresh L0 runs before and after in the same session.
- L2b confound: Cilium plus WireGuard. The L2b run added Cilium's chained eBPF datapath and WireGuard at the same time. A Cilium-without-encryption state was attempted but not benchmarked, because
helm upgradedid not restart the agents and WireGuard stayed active. The reported L2b difference is therefore Cilium chaining plus WireGuard combined; the WireGuard-only share was not isolated. - L2b drift. The L0-to-L0 drift within the L2b session was up to about 3.4% on throughput and 4.5% on E2E p50 (both at c=16). The c=1 and c=4 L2b difference is within that drift. With n=1 per run, even the c=16 and c=32 difference rests on one sample per level.
- L2b encryption proof is counter-based. The evidence is Cilium agent status, WireGuard peers and handshakes, and
cilium_wg0byte counters matching a known transfer. No packet capture showed ciphertext on the ENI. - L2b L0-after is not a proven Cilium-free baseline. 24-bracket-state shows the Cilium agent reinstalled and running with encryption off (no
cilium_wg0) during L0-after, and the final uninstall (30-cilium-uninstall-final) was at 13:22:22Z, after the sweep. The head pod was shown not to be a Cilium endpoint; the worker pod was not checked. If the worker was a Cilium endpoint, the bracket mean includes some Cilium datapath effect and understates the L2b difference. Two actions in the L2b session (the conflist rename at 13:09Z and the GPU scale-down from 13:22 to 13:24Z) were taken by the same study team in a parallel control session; neither overlaps a benchmark window. - No raw per-request samples.
bench/bench.pyline 119 writes{"summary", "raw"}inside the pod, butmanifests/stageB-bench-job.yamlline 49 prints only['summary'], and only the printed line was kept. No distributions, confidence intervals, bootstrap estimates, or outlier analysis are possible from the archive, and the cause of the multi-second stalls cannot be investigated after the fact. - Tag-pinned artifacts. Images were referenced by tag (for example
vllm/vllm-openai:v0.11.0), and runtime image digests were not recorded. The Linkerd version and the Checkov version were not recorded. - One run per level. Each layer and concurrency level has exactly one 128-request run, including all three L2b-session runs. There are no repeated trials, no confidence intervals, and no significance tests. The observed spread between runs (for example, throughput at c=32 from 429.2 to 568.8 tok/s across layers that should not affect throughput) indicates that run-to-run variance is large on this topology.
- Different pod instances. L1 ran on the L0 pods after one restart; L2a and L3 ran on newly created pods on the same nodes. Pod-level warm state (allocator state, Ray actor placement, page cache) differs between layers.
- Unexplained stalls. Multi-second TTFT outliers appear in L0 and L3 but not L1. Their cause is unknown, and they dominate tail and throughput comparisons.
- Percentile resolution. With 128 samples, p99 is set by one or two requests.
- Closed-loop client. Throughput is coupled to latency; a single stalled request lowers throughput for the level. An open-loop (fixed arrival rate) harness would separate the two.
- Startup window. A search of every stored log for
NETWORK_POLICY_ENFORCING_MODEfound the name only in an EKS add-on option listing, with no value. The logs show--enable-network-policy=trueon theaws-eks-nodeagentcontainer and nothing about the enforcing mode. Ray pods created while the policy already existed (for example after a restart) were never probed during startup. L1 restarted both engine pods, and the probe ran on the re-formed, long-running pods. Whether a newly created Ray pod is reachable from another namespace during its startup window in this cluster is an open question that this study cannot answer. - Proof coverage. The L1 probe cannot tell a dropped SYN from a non-listening port, so only ports open at L0 count as evidence. L1 was ingress-only and no egress test was run. All probes were steady-state; the VPC CNI
NETWORK_POLICY_ENFORCING_MODEvalue was not recorded (only the variable name appears in 01-vpccni-encryption-options), so a default-allow startup window understandardmode is unmeasured. The L3 proof covers ordinary requests to/v1/modelsand/v1/completionsonly; the CVE-2026-48746 Host-header bypass, which applies to the tested version, was not tested. - Ray native security not evaluated. Ray 2.52.0 supports token authentication and Ray supports TLS; both were off. The study does not compare network-layer controls against these application-layer controls.
External validity.
- Heterogeneous GPUs. PP stage 0 on an A10G and stage 1 on a T4 is not a production configuration. The T4 gates throughput, and the settings needed to make it work (
--enforce-eager, TORCH_SDPA, no FlashInfer sampler) change the performance profile. - Hardware change for L2b. The L2b session ran on 2 x A10G, not A10G + T4, with the same engine flags. c=32 throughput was about 1180 tok/s there against 429 tok/s on the Stage C pair. Absolute numbers are not comparable across the two sessions; only within-session deltas are. The percent difference may differ on the T4 pair, where the pipeline was slower and the network a smaller share of step time.
- Small model. Qwen2.5-1.5B in fp16 moves far fewer bytes per step between stages than a 70B-class model. An encryption layer's overhead is expected to scale with inter-node bytes, so the L2b difference here may understate the difference for larger models (inference, not measured).
- Topology. PP=2 over TCP, no EFA, no tensor parallelism across nodes, no disaggregated KV transfer (that path failed to start). The KV-transfer channel, which carries prompt-derived state, was not measured.
- One region, one day, on-demand fallback. All GPU runs were in ap-southeast-2 on 2026-09-28. Stage C/D ran on on-demand nodes because spot capacity was unavailable; results say nothing about spot interruption behaviour.
- Versions. vLLM v0.11.0, Ray 2.52.0, KubeRay 1.7.1, Linkerd edge (version not recorded), VPC CNI network policy agent. Defaults change across releases; for example, the Job API exposure depends on image contents, and the vLLM API-key behaviour differs between versions inside and outside the CVE-2026-48746 range.
- Short prompts. About 89 prompt tokens and 128 output tokens per request. Long-context or long-output workloads change the ratio of prefill to decode and of control traffic to data traffic.
Construct validity.
- "Plaintext" is shown for the Ray GCS and raylet listeners by the cleartext HTTP/2 preface and a failed TLS handshake. Encryption of the NCCL activation stream was not measured.
- "No measurable difference" means no difference larger than the variance we could observe with n=1. It is not evidence of zero effect.
- "L3 authenticated the engine" would be the wrong construct. What was measured is that ordinary requests on two
/v1routes were rejected without the key. - "L1 isolated the engine" would also be the wrong construct. What was measured is steady-state ingress blocking from one neighbour pod to specific ports.
Related Work
Citation confidence is marked. [C] means the source, its identifier and the detail cited were checked (all entries were re-checked on 2026-09-29).
Ray exposure. Ray's documentation states that Ray assumes a trusted network and that the dashboard and Job API must not be exposed to untrusted clients [1] [C]. CVE-2023-48022 describes unauthenticated remote code execution via the Ray Jobs API; Anyscale disputes it as a vulnerability on the grounds that Ray is designed to run inside a trusted boundary [2] [C]. Oligo Security reported in-the-wild exploitation of internet-exposed Ray clusters under the name ShadowRay in 2024 [3] [C]. Our contribution is the in-cluster version of the same exposure: not internet-facing, but reachable from other pods in a flat Kubernetes network unless a NetworkPolicy blocks it, and dependent on image contents. Ray token authentication exists from Ray 2.52.0 (RAY_AUTH_MODE=token), is disabled by default, and authenticates the dashboard, GCS and other control-plane services; KubeRay supports configuring it [4] [C]. Ray 2.52.0 was the version under test, and we ran it with token authentication off and did not evaluate it.
vLLM API authentication. CVE-2026-48746 / GHSA-94f4-hr76-p5j6 is a critical (CVSS 9.1) API-key bypass in the vLLM OpenAI API server: a crafted Host header makes AuthenticationMiddleware reconstruct a path that skips the /v1 check. It affects >=0.3.0, <0.22.0, is fixed in 0.22.0, was published 2026-06-02, and does not affect deployments behind an RFC-conforming proxy such as nginx [19] [C]. The tested v0.11.0 engine was exposed directly and is in range; we did not test the bypass.
Inference-engine advisories. vLLM has published advisories for unauthenticated deserialisation on its distributed communication paths, including the PyNcclPipe KV-transfer service (CVE-2025-47277) [5] [C] and ZeroMQ sockets in the multi-node V0 engine (CVE-2025-30165) [6] [C]. A 2025 report by Oligo Security, which named the pattern ShadowMQ, described the same unsafe ZeroMQ plus pickle pattern across several inference frameworks including vLLM and SGLang [7] [C]. These advisories concern code paths we did not exercise (the disaggregated connector failed to start, and we did not send data to any tensor channel). They support the general point that inter-node inference channels are an attack surface and not just a performance concern.
LLM serving systems and side channels. vLLM's PagedAttention design [8] [C] and Ray [9] [C] are the systems under test. Automatic prefix caching creates shared state across requests; timing side channels on shared KV and prefix caches in LLM serving have been studied [10] [C]. Our observation that /metrics exposes prefix-cache hit counters unauthenticated is a possible amplifier of such channels; we did not test it.
Kubernetes NetworkPolicy. The Kubernetes documentation defines NetworkPolicy semantics and states that enforcement depends on the network plugin [11] [C]. Budigiri et al. evaluated the performance and security of NetworkPolicy implementations across CNIs [12] [C]. Our L1 result, no measurable steady-state difference on the VPC CNI eBPF agent, is consistent in direction with the expectation that L3/L4 filtering is cheap, but n=1 does not allow a quantitative comparison.
Service mesh and encryption overhead. Zhu et al. dissected the latency and CPU overhead of sidecar service meshes and attributed much of it to L7 processing [13] [C]. Istio and Linkerd both publish performance measurements [14] [C]. WireGuard's design and performance are described by Donenfeld [15] [C], and Cilium documents transparent WireGuard encryption between nodes [16] [C]. We found no prior measurement of mesh or node encryption overhead specifically on Ray-backed multi-node LLM inference; L2b gives one bracketed data point for this testbed (about 3% to 8%, Cilium chaining plus WireGuard combined, n=1). The L2a failure is a configuration-specific negative result, not an overhead result and not a demonstrated incompatibility.
Static analysis of Kubernetes manifests. Trivy, Checkov, Kubescape, and kube-linter are widely used open-source configuration scanners [17] [C]. We are not aware of a published comparison focused on their handling of operator custom resources; that comparison would be useful and is not claimed here.
Ethics and Responsible Disclosure
- All testing ran in a dedicated Amazon EKS cluster and a dedicated VPC (
10.99.0.0/16) in AWS, created for this study and used for nothing else. Only our own pods were probed. - No third-party systems, no internet-facing services, and no other tenants were probed. Every probe target came from an allowlist of our own pod IPs; no CIDR ranges were scanned.
- Exploitation was limited to one benign synthetic Ray Job submission (marker print and hostname), which failed because the endpoint was not present. No state-changing vLLM routes were called. No tensor or KV data was injected, tampered with, or decoded. Prompts contained only synthetic text (a fake record number in the tokenizer test).
- The behaviours reported (Ray's trusted-network assumption, vLLM's optional authentication, charts shipping no NetworkPolicy) are documented vendor defaults rather than new vulnerabilities, so no new coordinated disclosure is required. The vLLM API-key bypass relevant to the tested version (CVE-2026-48746) was already public and fixed upstream. If the Linkerd failure is reproduced with a minimal, version-pinned configuration and a root cause, it will be reported to the relevant upstream issue tracker as a bug, not a security issue.
- Logs were redacted of names, emails, account and principal IDs, public IP addresses, and credentials before publication. GPU capacity was scaled to zero when the work finished.
Future work
Follow-ups drawn from the two independent reviews (Appendix D and Appendix E), ordered by how much they would change the conclusions.
- Route-by-route authentication matrix across vLLM versions. Enumerate every route from
/openapi.jsonand call each with no key, a wrong key, and the correct key, on v0.11.0, a patched release (0.22.0 or later), and each behind an RFC-conforming proxy. Reproduce the CVE-2026-48746 regression only in an isolated environment. Classify each route as inference-capable, administrative, or information-disclosing. - Ray native security versus network controls. A factorial of Ray token authentication (
RAY_AUTH_MODE=token) and Ray TLS, crossed with NetworkPolicy and WireGuard, measuring both security effect (rejection, handshake behaviour, reachability) and serving difference (throughput, TTFT, TPOT, E2E, CPU). - VPC CNI startup-window measurement. Probe a newly created target pod continuously from an unrelated namespace while recording sandbox creation, PolicyEndpoint programming, and Ready state; repeat many times under
NETWORK_POLICY_ENFORCING_MODE=standardand understrict, and report the exposure-window distribution. No GPU needed. - A/B/C Cilium decomposition. A: plain VPC CNI. B: Cilium chained with encryption disabled (with agents restarted and the state confirmed on every engine pod). C: the same Cilium plus WireGuard. Randomise the order across at least 10 blocks on identical nodes so that B minus A estimates the datapath difference and C minus B estimates the incremental WireGuard difference, with confidence intervals over blocks.
- Egress tests. Add egress NetworkPolicy and test outbound reach from a compromised-engine vantage point (other namespaces, instance metadata, the internet, model stores).
- Digest-pinned artifact manifest. Record runtime image digests, chart digests, model and tokenizer revisions, scanner, Linkerd, Cilium, kernel, driver and CNI versions in one machine-readable manifest with SHA-256 sums, and keep the harness under version control.
- Persist raw samples. Change the bench Job so the full
{"summary","raw"}file leaves the pod (print it, or write it to a volume), so every run keeps per-request TTFT, TPOT and E2E for distributions, bootstrap intervals, and stall analysis. - Linkerd root cause. Repeat L2a with a pinned Linkerd version, kept proxy logs, and a controlled split of Service versus pod-IP targeting, protocol detection versus opaque ports, and one-sided versus two-sided meshing. This is an engineering question, not a security one.
- What the exposed Ray ports allow beyond reachability. Not tested: whether a client with network access to GCS and raylet alone can read cluster state, schedule work or execute code on the GPU inference deployment. This needs an approved, isolated experiment.
- Ray native authentication and NetworkPolicy together. Not tested: whether NetworkPolicy is still needed when Ray token authentication or TLS is configured, and what each closes on its own (see item 2).
- Scanner fairness. Not tested: the same manifests with custom rules, with admission-time policy engines that render custom resources, and with scanners that parse the RayCluster resource, so that a like-for-like comparison of runtime-exposure detection is possible. First step: render the RayCluster into its resulting head and worker pod specs and rerun the same four scanners on them. The comparison in Four default scanner configurations and the runtime surface is not like for like.
- Newly created Ray pods during startup. Not tested (see item 3): probe pods created after the policy exists, with the enforcing mode recorded first.
- Job API and code execution on the GPU stack. Not demonstrated. If the Job API is enabled in a GPU deployment, test it only in an isolated environment with the same safe-probing rules.
Conclusion
In the tested KubeRay and vLLM topology, distributed inference introduced Ray control-plane listeners that were not represented by the workload's declared port metadata, and an ordinary pod in an unrelated namespace could reach them: Ray GCS and raylet spoke cleartext gRPC on ports that are mostly absent from the pod's declared port metadata. This is one cluster, one topology and one version set, and the reach was shown from one constructed neighbour pod. An ingress-only default-deny NetworkPolicy blocked cross-namespace ingress from that neighbour to the Ray control-plane ports in steady state; it gives no egress containment, and the pod startup window was not tested. The Ray Job API was exposed on the default CPU Ray image and was not listening in the GPU vLLM deployment, so remote code execution against the GPU inference stack was not demonstrated. Four default scanner configurations did not inspect the RayCluster custom resource, so the generated Ray pods were not represented in their analysis and the runtime network exposure was not identified; custom rules were not tested, and these are runtime and network properties that a reachability test measures. The weaker results follow. An engine API key turned ordinary unauthenticated requests on two /v1 routes into 401, which shows the normal key path works and does not show that API-key authentication protects vLLM: the other routes were not tested and the tested v0.11.0 was, at experiment time, publicly known to be bypassable through CVE-2026-48746 (not tested here). In a single bracketed run, Cilium chaining plus WireGuard combined carried every streaming request and routed inter-node pod traffic through the tunnel, with about 3% to 8% lower throughput and higher median latency, the low-concurrency part inside session drift and the closing bracket not proven Cilium-free; WireGuard alone was not isolated and NCCL traffic was not separately captured. Ray's own token authentication and TLS, available in the tested version, were not evaluated and are the obvious next comparison. Two engineering incidents are recorded for completeness: the tested Linkerd configuration failed to carry the vLLM workload (root cause not isolated), and an incomplete Cilium removal procedure broke new pods on every node until a stale CNI config was cleaned up. The single most important methodological lesson is that performance comparisons on a two-GPU pipeline need same-session baselines, repeated trials, and persisted per-request samples; only the L2b session had the first, and none of the runs had the other two.
Sorami's view
This section is opinion. It is how we would prioritise the fixes for a team running distributed inference.
Treat the Ray network as internal plumbing that must never be reachable from other workloads. Apply an ingress default-deny NetworkPolicy to the Ray namespace first, because it closed the most for the least effort in our test. Then add node-level encryption and check Ray's own authentication. Upgrade vLLM past 0.22.0 before trusting its API key, which is a weak layer on its own.
The four default scanner configurations we ran did not report this surface, and reachability is a runtime property. A reachability test from a neighbour pod measures it. That is part of our AI production readiness review and our cloud penetration testing.
What to do now
- Apply an ingress default-deny NetworkPolicy to every Ray namespace.
- Allow only the router to reach the vLLM API port.
- Evaluate Ray token authentication, which is off by default.
- Probe your Ray ports from a pod in another namespace.
- Encrypt inter-node pod traffic with WireGuard or IPsec.
- Upgrade vLLM to 0.22.0 or later, then set an API key.
Our AI on Kubernetes Helm chart research covers the single-chart defaults behind this stack. The practical guide gives the steps in order, and the guides index has related material.
Frequently asked questions
Is Ray secure to run on a shared Kubernetes cluster?
Not with defaults. The Ray documentation says a Ray cluster should run only inside a trusted network. In our one-cluster test the Ray control ports accepted cleartext gRPC from a pod in an unrelated namespace. Ray token authentication exists from Ray 2.52.0 but is off by default and was not evaluated here.
Do Kubernetes security scanners detect exposed Ray ports?
Not in the default configurations we tested. Trivy, Checkov, Kubescape and kube-linter did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis. Reachability is a runtime property, and custom rules were not tested. 15 of 17 observed Ray listening sockets were not declared as container ports, which are metadata only, so a review that relies on them has incomplete visibility.
Does a vLLM API key protect the server?
Not on its own. With an API key set, ordinary requests to /v1/models and /v1/completions got 401 instead of 200, which shows the normal key path works and does not show the API is protected. Other routes were not tested. The tested version is in the affected range of CVE-2026-48746, a Host header bypass fixed in vLLM 0.22.0, so upgrade before relying on the key.
How much did Cilium with WireGuard change LLM inference speed?
In one same-session bracket, Cilium chaining plus WireGuard combined showed 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median full-request latency. The difference rose with concurrency. It does not isolate WireGuard, and each figure is a single run on a small 1.5B model.
What is the cheapest fix for an exposed Ray cluster?
An ingress default-deny NetworkPolicy. In our test it blocked every Ray port that was open to the neighbour pod in steady state. The startup window for newly created pods was not tested, and the performance comparison is a single run per level.
Appendix: References
All identifiers and URLs below were checked on 2026-09-29.
- Ray Project. "Security" and "Configuring and managing the Ray Dashboard" in the Ray documentation. [C]
- NIST National Vulnerability Database. CVE-2023-48022 (Ray Jobs API, disputed by vendor). https://nvd.nist.gov/vuln/detail/CVE-2023-48022 [C]
- A. Lumelsky et al., Oligo Security. "ShadowRay: First Known Attack Campaign Targeting AI Workloads Actively Exploited In The Wild", 26 March 2024. Verified 2026-09-29. [C]
- Ray Project. "Ray token authentication" (available from Ray 2.52.0,
RAY_AUTH_MODE=token, disabled by default); and "KubeRay authentication", https://docs.ray.io/en/releases-2.52.0/cluster/kubernetes/user-guides/kuberay-auth.html. Verified 2026-09-29. [C] - vLLM Project. GHSA-hjq4-87xh-g4fv / CVE-2025-47277, "Remote Code Execution via PyNcclPipe Communication Service" (V0 engine
PyNcclPipeKV-transfer only, affected >=0.6.5 <0.8.5, fixed 0.8.5), published 20 May 2025. Verified 2026-09-29. [C] - vLLM Project. GHSA-9pcc-gvx5-r5wm / CVE-2025-30165, "Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration" (ZeroMQ SUB socket plus pickle in the multi-node V0 engine; not fixed, V0 off by default since 0.8.0), published May 2025. Verified 2026-09-29. [C]
- A. Lumelsky, Oligo Security. "ShadowMQ: How Code Reuse Spread Critical Vulnerabilities Across the AI Ecosystem", November 2025 (unsafe ZeroMQ
recv_pyobj/pickle pattern in Meta Llama Stack, NVIDIA TensorRT-LLM, vLLM, SGLang, Modular Max Server and others). Verified 2026-09-29. [C] - W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, I. Stoica. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023. [C]
- P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, I. Stoica. "Ray: A Distributed Framework for Emerging AI Applications." OSDI 2018. [C]
- X. Zheng et al. "InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks". arXiv:2411.18191, November 2024; and L. Song, Z. Pang, W. Wang et al. "The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems." arXiv:2409.20002, September 2024, https://arxiv.org/abs/2409.20002 (published in IEEE Transactions on Information Forensics and Security, vol. 20, 2025). Verified 2026-09-29. [C]
- Kubernetes documentation. "Network Policies". [C]
- G. Budigiri, C. Baumann, J. T. Muehlberg, E. Truyen, W. Joosen. "Network Policies in Kubernetes: Performance Evaluation and Security Analysis". EuCNC/6G Summit 2021, pp. 407-412. Verified 2026-09-29. [C]
- X. Zhu, G. She, B. Xue, Y. Zhang, Y. Zhang, X. K. Zou, X. Duan, P. He, A. Krishnamurthy, M. Lentz, D. Zhuo, R. Mahajan. "Dissecting Overheads of Service Mesh Sidecars". ACM SoCC 2023, pp. 142-157. Verified 2026-09-29. [C]
- Istio documentation, "Performance and Scalability"; Linkerd, published benchmark posts. [C]
- J. A. Donenfeld. "WireGuard: Next Generation Kernel Network Tunnel." NDSS 2017. [C]
- Cilium documentation. "WireGuard Transparent Encryption". [C]
- Project pages: Trivy https://github.com/aquasecurity/trivy, Checkov https://github.com/bridgecrewio/checkov, Kubescape https://github.com/kubescape/kubescape, kube-linter https://github.com/stackrox/kube-linter [C]
- Amazon EKS User Guide. "Configure network policy" (
NETWORK_POLICY_ENFORCING_MODE: defaultstandardstarts new pods default-allow until policies are programmed;strictstarts default-deny). Verified 2026-09-29. [C] - GitHub Advisory Database. GHSA-94f4-hr76-p5j6 / CVE-2026-48746, vLLM OpenAI API server authentication bypass via crafted Host header, CVSS 9.1, affected >=0.3.0 <0.22.0, fixed 0.22.0, published 2026-06-02. https://github.com/advisories/GHSA-94f4-hr76-p5j6; OSV https://osv.dev/vulnerability/CVE-2026-48746. Verified 2026-09-29. [C]
- Cilium documentation. "Helm Reference" (
cni.uninstall, defaultfalse). The default is documented; the URL was not re-fetched on 2026-09-29. [C]
Appendix A: Reproduction
The harness is in the test-harness/ folder of this repository. Paths below are relative to it. Copy infra/terraform.tfvars.example to infra/terraform.tfvars and set operator_cidr to your own /32. Every step that mutates infrastructure must be scoped to a dedicated, isolated account or VPC; do not run the neighbour probe in a shared cluster.
- Infrastructure.
infra/(Terraform): EKS 1.35 with VPC CNIenableNetworkPolicy=true, a CPU spot group, a GPU spot group (taintedsorami-lab/gpu=true:NoSchedule), and a GPU on-demand fallback group (max 2 for Stage C). Scale GPU groups with-var gpu_desired=N -var gpu_ondemand_desired=M. - GPU device plugin. NVIDIA device plugin Helm release restricted to the GPU pool.
- Stage A. Install KubeRay operator 1.7.1,
ray-cluster1.7.1, and vLLM production-stack 0.1.12 with default values. Render them tomanifests/rendered-defaults/all-defaults.yamland runmanifests/scanner-jobs.yaml. - Neighbour probe.
manifests/neighbor-probe.yaml: namespacesorami-lab-neighbor, uid 1000, all capabilities dropped, no ServiceAccount token. Build the allowlist from the endpoints list, then runnmap -sT -Pn -p- --openagainst only those IPs andscripts/handshake.shfor read-only handshakes. - Stage B.
manifests/stageB-vllm-engine.yaml(Qwen2.5-7B-Instruct,vllm/vllm-openai:v0.11.0), thenscripts/stageB-bench-run.shandscripts/stageB-reprobe.sh. - Stage C.
manifests/stageC-ray-pp.yamlwith node names pinned to your two GPU nodes. Benchmark:scripts/stageC-bench-run.sh <logdir> "<label>". - Stage D. L1:
kubectl apply -f manifests/stageD-L1-networkpolicy.yaml, then re-probe and benchmark. L2a: install Linkerd, annotate the engine pods (skip Ray ports 6379 and 10002 to 10010), run the bench withMESH_BENCH=1. L2b: remove the L1 policies, run a fresh L0,helm install cilium cilium/cilium --version 1.20.2 -n kube-system -f manifests/stageD-L2b-cilium-values.yaml(setcni.uninstall=truefrom the start), recreate the engine pods, confirmcilium-dbg statusshows WireGuard, run the bench, then remove Cilium, confirm no05-cilium.conflistremains on any node, recreate the engine pods, and run L0 again. L3: create Secretsorami-lab-vllm-apikey(keyapikey), restart the RayCluster, run the bench withBENCH_AUTH=1. - Method fixes this study did not have: re-run L0 immediately before each layer in the same session, repeat each level in at least 10 randomised blocks, alternate layer order (for example L0, L1, L0, L3, L0), persist the per-request
rawlist thatbench/bench.pyalready writes inside the pod (the Job atmanifests/stageB-bench-job.yamlline 49 currently prints onlysummary), and report medians with bootstrap confidence intervals over blocks. Record image digests, not only tags. Consider an open-loop arrival-rate client to decouple throughput from stalls. - Teardown. Scale both GPU groups to 0 and delete the
sorami-lab-*namespaces;terraform destroyremoves the cluster and VPC.
Appendix B: Artifact list
Findings summaries: stageA-baseline, stageB-baseline, stageC-distributed, stageD-hardening-tax. Repository guide: README.
Appendix C: Data consistency notes
Discrepancies found while checking the findings summaries against the raw logs. The report uses the raw values.
| Item | Summary said | Raw log shows | Effect |
|---|---|---|---|
| Stage D L0 | "re-confirm the distributed baseline" (study plan); L0 listed as "measured" | Byte-identical copy of Stage C results, no run logs | Disclosed in Baseline provenance (important) and Limitations |
| L1 at c=32 | Findings table showed +32.5% throughput | Correct arithmetic, but attributable to the L0 stall, not to L1 | Reframed as no measurable difference |
| Stage C worker ports | Findings table lists only 10002 on worker as open | Matches; the L1 probe also marks non-listening worker ports as BLOCKED | Only ports open at L0 counted as L1 evidence |
| L2a proxy behaviour | Findings describe L7 503 and opaque-mode reset | Result files contain only the 503 run; opaque-mode behaviour is in the outcome note, not in captured proxy logs | Opaque-mode claim labelled as operator note |
| L3 auth proof | 4 calls listed | Log has 6 lines, including one correct-key completion that timed out (curl exit 28) before a successful retry | Reported both |
Stage B /metrics families | 66 | 12-metrics-names.log counts 66 # HELP lines | Consistent |
| Stage A findings "7 ephemeral RPC ports" on head | 7 | nmap shows 7 high ports (36087, 36349, 44217, 44227, 52365, 57914, 62226) | Consistent |
| Checkov version | not stated | not logged | Stated as unknown |
| Linkerd version | "edge" | exact version not logged | Stated as unknown |
| L2b comparison basis | Findings first reported L2b vs L0-before only (-6.2% c=16, -5.3% c=32 throughput) | Both bases are correct arithmetic from the JSON; vs mean of both L0 runs gives -4.5% and -6.5% | The report uses the mean of both L0 runs; L0-before-only values are in 50-wireguard-tax.md |
| L2b range | Findings said "about 4 to 6% throughput, 4 to 8% e2e" | vs L0 mean: 3.4% to 6.5% throughput, 3.2% to 8.4% E2E p50 | The report gives 3% to 8% overall with per-level values |
| L2b L0-after state | Outcome note says pods on plain VPC CNI path | Head pod confirmed not a Cilium endpoint; worker not checked | Stated as head-only evidence |
| Stage C GPU peaks | "head ~13%, worker ~68%" | head max 15%, median 12%; worker max 68%, median 37% | The report gives min, median, max |
| Raw per-request data (added 2026-09-29) | Tables implied distribution-level analysis was possible | bench/bench.py line 119 writes {"summary","raw"} inside the pod; manifests/stageB-bench-job.yaml line 49 prints only ['summary']; every archived result-c*.json has only p50/p95/p99/mean | Stated in Benchmark harness and Limitations; no distributions or CIs reported |
| L1 proof port count (added 2026-09-29) | Summary of layers said "6 open ports to blocked" | 7 ports open at L0 went to BLOCKED: head 6379 and 10002 to 10006 (6), worker 10002 (1) | Corrected to 7 |
| L1 scope (added 2026-09-29) | "L1 closes the cross-namespace exposure" | Every policy is policyTypes: ["Ingress"] (line 22 for default-deny); egress deliberately open (comment at line 11) | Stated as ingress-only, no containment claim |
| L3 scope (added 2026-09-29) | "removed unauthenticated inference" | 01-auth-proof.log covers /v1/models and /v1/completions only; v0.11.0 is in the CVE-2026-48746 range | Narrowed to two /v1 routes; CVE added |
| L2a wording (added 2026-09-29) | "incompatible" | Only the 503 result files exist; no proxy logs, no Linkerd version | Reworded to "tested configuration failed; root cause not isolated" |
| L0-after bracket (added 2026-09-29) | "valid no-encryption L0", "plain VPC CNI path" | 24-bracket-state.log: Cilium agent running, encryption off, no cilium_wg0; head not an endpoint; worker unchecked; final uninstall at 13:22:22Z | Stated as not a proven Cilium-free baseline |
| Actor for rename and scale-down (added 2026-09-29) | attributed to a different session actor | Same study team, parallel control session | Wording corrected here and in 60-encryption-proof-and-outcome.md |
| VPC CNI enforcing mode (added 2026-09-29) | not discussed | 01-vpccni-encryption-options.log lists NETWORK_POLICY_ENFORCING_MODE as a variable name only, no value | Startup window stated as unmeasured |
Appendix D: Response to independent review (2026-09-29)
Corrections after two independent reviews were applied on 2026-09-29. This appendix and Appendix E record them.
Two independent reviews of the draft were received on 2026-09-29. One review was received as a secondary copy whose size and hash did not match what the primary review described, so we treated it as a secondary copy of the review points and checked every point against the raw logs and harness ourselves. External facts were re-verified on the web on 2026-09-29 (references [4], [18], [19]).
| Review point | Decision | Where fixed |
|---|---|---|
API key "removed unauthenticated inference" is too strong; only two /v1 routes were tested | Accepted | Summary, Contribution 4, Reachability matrix, Security effect of each layer, Summary of layers, Practical recommendations, Defence in depth ordering, Limitations, Conclusion |
| vLLM v0.11.0 is in the CVE-2026-48746 / GHSA-94f4-hr76-p5j6 affected range | Accepted (verified: >=0.3.0 <0.22.0, published 2026-06-02; bypass not tested here) | Distributed LLM inference, Software under test, Security effect of each layer, Related Work, References [19] |
| Ray token authentication and TLS were omitted | Accepted (verified: token auth from 2.52.0, off by default; not evaluated). Reference [4] upgraded from [U] to [C] | Kubernetes networking and NetworkPolicy, Software under test, Practical recommendations, Defence in depth ordering, Limitations, Related Work, Future work |
| WireGuard figure is Cilium chaining plus WireGuard, n=1 | Accepted; already disclosed, wording tightened to "associated with" | Summary, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Network-layer versus L7 encryption for streaming inference, Conclusion |
| Linkerd "incompatible" is not demonstrated | Accepted | Contribution 5, L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload), Summary of layers, Network-layer versus L7 encryption for streaming inference, Related Work, Ethics and Responsible Disclosure, Conclusion |
| Per-request raw samples were never persisted | Accepted (verified: bench/bench.py line 119 vs manifests/stageB-bench-job.yaml line 49) | Benchmark harness, Limitations, Appendix A, Appendix C |
| L1 is ingress-only; no containment evidence | Accepted (verified: line 22 policyTypes: ["Ingress"], line 11 egress left open) | Kubernetes networking and NetworkPolicy, Hardening layers, Security effect of each layer, Summary of layers, Practical recommendations, Defence in depth ordering, Limitations |
containerPort is metadata, not a firewall | Accepted | Summary, Contribution 2, Kubernetes networking and NetworkPolicy, Four default scanner configurations and the runtime surface, What the four default scanner configurations did not report, and why |
VPC CNI standard mode has a default-allow startup window | Partly accepted: the documented behaviour is real, but no evidence either way for this study, because the value was not recorded. Presented as a gap, not a finding | Kubernetes networking and NetworkPolicy, Environment, Reachability matrix, Security effect of each layer, Limitations |
| Cilium removal failure should be framed as incomplete procedure with a documented default | Accepted | Summary, Contribution 7, L2b: Cilium chaining plus WireGuard combined (single bracketed run) |
| L0 reused for L1 and L3 is cross-session | Accepted; already disclosed in 4.6 | no change |
| Tag-based images; digests and some versions not recorded | Accepted | Software under test, Limitations, Future work |
| Wilson lower bound for 128 of 128 successes is about 97% | Accepted as a note; not added to tables, since success counts are not a headline claim | none |
Proposed designs: 10 blocks, 256 measured requests per run, randomisation seed 20260929 | Partly accepted: the 10-block minimum is adopted in Future work; the specific request count and seed are the reviewers' proposal, not requirements | Future work |
| Cost per million tokens and a Pareto frontier | Partly accepted as a later extension; not in this report's scope | not added |
| Infrastructure-agent misalignment study and telemetry-assisted prefix-cache leakage | Rejected for this report: outside the research questions. May be separate studies | not added |
| "Universal cluster-wide reachability" should not be claimed | Accepted; the report claims reach from one neighbour pod in one testbed | Conclusion |
Items from the reviews that were already correct in the report (L0 reuse, counter-based rather than pcap WireGuard evidence, L2b drift) were left as they were.
Appendix E: Response to a second independent review (2026-09-29)
A second independent review was received on 2026-09-29. It rated the cross-namespace Ray exposure, the cleartext Ray control-plane evidence and the NetworkPolicy blocking as the strong results, and the performance numbers and generalisation as the weakest. We checked each point against the report and the stored logs. No experiment was run and no measured number was changed. Only wording, framing and structure changed, plus an offline re-reading of the stored scanner outputs.
| Review point | Decision | What changed | Where |
|---|---|---|---|
| Headline overstated: one constructed neighbour pod, one cluster | Accepted. The old conclusion generalised the result to the whole cluster, which the test did not show | Title, subtitle, abstract and conclusion now state the constrained result: a pod in an unrelated namespace could reach Ray control-plane listeners under these defaults, in one cluster, one topology and one version set | Title, Summary, Conclusion |
| Scanners: the claim could be read as a general one | Accepted | Restated as four default scanner configurations that did not identify the RayCluster's runtime network exposure. Four default scanner configurations and the runtime surface says these are runtime and network properties and that custom rules were not tested. Offline re-analysis of the stored outputs: no output names the RayCluster, and rules of the kind that fired on the two Deployments would plausibly apply to its pod templates. No rule concerns reachability. The 57-versus-zero comparison is not like for like | Summary, Introduction, Four default scanner configurations and the runtime surface, What the four default scanner configurations did not report, and why, Future work |
| containerPort argument is dramatic | Accepted | Reframed as visibility drift for tooling that treats declared ports as an exposure inventory. containerPort is metadata, not a security boundary | Summary, Introduction, Four default scanner configurations and the runtime surface |
| Performance called a "cost" from n=1 | Accepted. The numbers are unchanged | "Cost" and "tax" became "observed difference in this single bracketed run" in prose, headings and table headers. The limits are stated next to the figures | Summary, Introduction, Performance comparison, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion, Limitations, Conclusion |
| WireGuard confound | Accepted; the confound was already disclosed | The configuration is labelled Cilium chaining plus WireGuard combined throughout, and the text says WireGuard alone was not isolated. The A, B and C design was already in Future work | Summary, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion, Future work, Conclusion |
| API key result given too much weight | Accepted; the caveats were already present | Moved from a headline result to a weak one. It shows the normal key path returned 401 and does not show that API-key authentication protects vLLM. Recommendation order changed | Summary, scope box, Security effect of each layer, Practical recommendations, Defence in depth ordering, Conclusion |
| Ray RCE could be misread | Accepted | Added a claim-by-claim table. Job API can be exposed: yes. Default CPU Ray exposed it: yes. GPU vLLM exposed it: no. Remote code execution against the GPU inference stack: no. Ray control ports reachable cross-namespace: yes | Scope box, Summary, The distributed control plane: cleartext and reachable |
| NCCL encryption stated as fact | Accepted | Worded as inference from the WireGuard datapath carrying inter-node pod traffic. NCCL traffic was not separately identified or captured | Summary, The distributed control plane: cleartext and reachable, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion |
| Startup window unanswered | Accepted. A search of all stored logs found the variable name only in an add-on option listing, with no value | Recorded as not recorded and as an open question: newly created Ray pods during startup were not tested | Limitations, Future work |
| Linkerd is an incident, not a security finding | Accepted | Labelled an engineering incident and moved out of the main conclusions and takeaways. The data stays in L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload) | Summary, Introduction, L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload), Conclusion |
Left as it was: the L0 reuse disclosure, the counter-based encryption evidence, the drift analysis and all measured values. The reviewer's three open questions (what the exposed Ray ports allow beyond reachability, whether NetworkPolicy is still needed with Ray native authentication or TLS, and whether the scanner comparison is fair) would need new experiments or new scans. They are recorded in Future work as not tested.
Third round of wording changes (2026-09-29, no experiment, no measured number changed): generalisations beyond the tested KubeRay and vLLM topology were removed, the scanner result now leads with the RayCluster not being inspected (finding counts stay in the Four default scanner configurations and the runtime surface table), the Ray Job API recommendation now reads as a potential arbitrary-code-execution interface resting on documented semantics, a Demonstrated / Not demonstrated / Engineering incidents summary was added after the scope box, and Future work now names rendering the RayCluster into pods and rerunning the same scanners.
About Sorami
Sorami is an Australian cyber security and cloud consultancy. We build and secure cloud environments, Kubernetes and AI agents, and the senior engineers who scope the work deliver it themselves. To discuss this research or a review of your own stack, use the form below. Only your email is required, and we reply within one business day.