Home / Research / Distributed LLM inference

The Hidden Network: Ray Control-Plane Exposure in Distributed LLM Inference on Kubernetes

What a pod in an unrelated namespace could reach on default Ray and vLLM deployments in one Amazon EKS testbed, what four default scanner configurations reported, and what an ingress NetworkPolicy blocked.

Sorami Technical Report · Distributed LLM inference security
Experiment date: 28 September 2026 (UTC)
Published: 29 September 2026
Author: Sorami Consulting
Data and evidence: Sorami-Consulting-AU/distributed-llm-inference-hidden-network
Charts: evidence/summary.html
Suggested citation: Sorami (2026). The Hidden Network: Ray Control-Plane Exposure in Distributed LLM Inference on Kubernetes. Sorami Technical Report. https://sorami.com.au/research/distributed-llm-inference-hidden-network/

In one Amazon EKS cluster, a pod in an unrelated namespace could reach default Ray control-plane listeners, which spoke cleartext gRPC. An ingress NetworkPolicy blocked those ports in steady state. Four default scanner configurations did not inspect the RayCluster resource. Sorami measured this once, on one topology.

Key takeaways

  • Ray GCS and raylet ports answered cleartext gRPC across namespaces.
  • An ingress NetworkPolicy blocked every Ray port that was open.
  • Default scanners did not inspect the RayCluster, so its Ray pods were not analysed.
  • 15 of 17 Ray sockets were missing from declared ports, which are only metadata.
  • Cilium chaining plus WireGuard differed by 3% to 8% in one run.
  • Job API code execution was not shown against the GPU inference stack.

Want the steps without the method? Our practical guide, how to secure Ray and vLLM on Kubernetes, turns these findings into a checklist. The full report follows.

Summary

Production LLM serving is moving from single-GPU engines to multi-node topologies built on Ray, vLLM, and KubeRay. In the tested KubeRay and vLLM topology, distributed inference introduced Ray control-plane and tensor-transfer listeners that were not represented by the workload's declared port metadata. We measured this surface on one dedicated Amazon EKS 1.35 cluster running vendor-default charts and images, from one constructed neighbour pod in an unrelated namespace.

Under these defaults, in this cluster, topology and version set, that pod could reach Ray control-plane listeners: a two-node Ray pipeline-parallel vLLM deployment exposed Ray GCS (6379) and raylet RPC (10002 to 10006) cross-namespace, and they spoke cleartext gRPC. An ingress-only default-deny NetworkPolicy (L1) turned every probed Ray port that was open at L0 into BLOCKED from the neighbour pod in steady state; the pod startup window was not tested.

The Ray Job API (8265) answered unauthenticated requests on the default CPU Ray image but was not listening in the GPU vLLM deployment, so remote code execution against the GPU inference stack was not demonstrated (claim table in The distributed control plane: cleartext and reachable). Four default scanner configurations (Trivy, Checkov, Kubescape, kube-linter) did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis, and they did not identify the RayCluster's runtime network exposure. These are runtime and network properties, and custom rules were not tested (Four default scanner configurations and the runtime surface). 15 of 17 observed Ray listening sockets were absent from the pods' declared port metadata: visibility drift for tooling that treats declared ports as an exposure inventory (containerPort is metadata, not a security boundary). A single-GPU vLLM engine exposed 26 routes with no authentication configured, including tokenizer and prefix-cache telemetry.

Two hardening results are weaker. An engine API key (L3) turned ordinary unauthenticated requests on two /v1 routes (/v1/models, /v1/completions) from 200 into 401, which shows the normal key path works and does not show that API-key authentication protects vLLM: the other routes were not tested, and the tested engine, vllm/vllm-openai:v0.11.0, is in the affected range of CVE-2026-48746 (a Host-header API-key bypass fixed in 0.22.0), which we did not test. The layers L1 and L3 showed no measurable steady-state difference within cross-session variance (n=1 run per level, median full-request (E2E) latency within 5% of baseline). Cilium chained on the VPC CNI with WireGuard (L2b, the combined configuration, not WireGuard alone) served 512 of 512 streaming requests, and interface counters showed inter-node pod traffic routed through the WireGuard device. NCCL traffic was not separately identified or captured, so NCCL encryption is inferred from that datapath and not measured. Against a same-session L0 bracket on a different GPU pair, the encrypted configuration produced 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median full-request (E2E) latency in this single bracketed run, rising with concurrency. The low-concurrency part of that difference is within L0-to-L0 drift, and the closing L0 bracket was not a proven Cilium-free state.

Separately, an engineering incident: the tested Linkerd mTLS sidecar configuration (L2a) failed to carry the vLLM workload (128 of 128 requests failed with HTTP 503 at every concurrency level), the root cause was not isolated, and it is not a security finding. An incomplete Cilium removal procedure (uninstall with the documented chart default cni.uninstall=false) left a stale CNI config on every node and broke new pod sandboxes for about 28 minutes. Results are limited by single runs, no persisted per-request samples, heterogeneous GPUs, a small model, and a reused baseline for L1 and L3, all disclosed.

What this study shows and does not show

Scope, environment and evaluation bounds:

  • Threat model: This study evaluates what an attacker who controls one ordinary pod in an unrelated namespace of the same cluster can reach. Container escape, node compromise, attacks on the Kubernetes API server, cloud IAM, supply chain, model-level attacks and attacks from outside the VPC are out of scope.
  • Environment: One dedicated Amazon EKS 1.35 cluster in ap-southeast-2, one day (2026-09-28 UTC), on-demand GPU nodes for the distributed runs, a small model (Qwen2.5-1.5B in fp16) and short prompts. Results say nothing about spot interruption, larger models or other regions.
  • Evaluation bounds: Each layer and concurrency level has one 128-request run, with no repeated trials, confidence intervals or significance tests. The Stage D L0 baseline is a reused copy of the Stage C run, so the L1 and L3 deltas are cross-session.
  • What the hardening results do not show: The API key result (L3) shows that the normal key path returned 401. It does not show that API-key authentication protects vLLM: only two /v1 routes were tested, and the tested engine is in the affected range of CVE-2026-48746, which we did not test. The encryption result is Cilium chaining plus WireGuard combined, not WireGuard alone, and it is one bracketed run. Ray token authentication and TLS were not evaluated.
  • What was and was not demonstrated: Ray control ports were reachable across namespaces, and the Ray Job API was reachable on the default CPU Ray image. The Job API was not listening in the GPU vLLM deployment, and remote code execution against the GPU inference stack was not demonstrated. The distributed control plane: cleartext and reachable has a claim-by-claim table.
  • Scanners: The result is about four default scanner configurations. Custom rules, admission-time policy engines and other scanners were not tested.

Demonstrated, not demonstrated and engineering incidents

Demonstrated in this testbed:

Not demonstrated:

Engineering incidents, not security findings:

Introduction

Inference engines such as vLLM expose an OpenAI-compatible HTTP API and, when scaled past one GPU, delegate placement and coordination to Ray. The Kubernetes operator KubeRay packages Ray as a custom resource. In the deployment we tested there were two sets of listeners: the declared one (a Service on port 8000) and a second set that the manifests did not declare (Ray GCS, raylet, object manager, worker RPC, and the collective-communication channel that carries activations or KV cache between GPUs). Ray's own documentation states that a Ray cluster should run only inside a trusted network. Kubernetes allows pod-to-pod traffic unless a NetworkPolicy restricts it. This study measures what that combination allowed in one testbed.

This report asks three questions:

  • RQ1: What does an unprivileged neighbour pod see when single-node and multi-node LLM inference stacks are deployed with vendor defaults?
  • RQ2: Do four widely used static Kubernetes scanners, in their default configurations, identify that runtime network exposure?
  • RQ3: What does each hardening layer close, and what difference in throughput and latency did we observe on the same topology?

Contributions:

  1. An empirical attack-surface inventory of vendor-default KubeRay 1.7.1, Ray 2.52.0, vLLM production-stack 0.1.12, and vllm/vllm-openai:v0.11.0 on EKS, from an unprivileged cross-namespace pod (Results: attack surface).
  2. Evidence that four default scanner configurations did not inspect the RayCluster custom resource, so the generated Ray pods were not represented in their analysis and the RayCluster's runtime network exposure was not identified (no scanner output contains a reference to the RayCluster resource), and that 15 of 17 observed Ray listening sockets (10 of 11 on the head, 5 of 6 on the worker) were absent from the pod's declared port metadata (containerPort). This is visibility drift for tooling that treats declared ports as an exposure inventory. containerPort is metadata, not a security boundary, and Kubernetes does not block undeclared ports (Four default scanner configurations and the runtime surface).
  3. Wire-level evidence that the multi-node Ray control plane is cleartext gRPC and reachable cross-namespace, and a claim-by-claim statement of what was and was not shown about the Ray Job API and remote code execution: exposed on the default CPU Ray image, not listening in the GPU vLLM deployment, and no code execution demonstrated against the GPU inference stack (The distributed control plane: cleartext and reachable).
  4. A layered hardening evaluation (L0 to L3) on a fixed two-node pipeline, with a per-layer security observation scoped to what was probed, and a performance comparison that is explicit about its statistical limits (Results: hardening). The L3 API key shows that the normal key path returned 401 on two /v1 routes; it does not show that API-key authentication protects vLLM, and the tested vLLM version was publicly known to be bypassable at experiment time (CVE-2026-48746, not tested here).
  5. A same-session, bracketed comparison of Cilium chained on the EKS VPC CNI with WireGuard (the combined configuration, not WireGuard alone) on Ray pipeline-parallel inference: it carried SSE streaming, and in this single bracketed run it produced about 3% to 8% lower throughput and higher median E2E latency, with the drift, the Cilium-plus-WireGuard confound, and the unproven Cilium-free state of the closing bracket stated (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
  6. An operational lifecycle hazard for Cilium chaining on EKS: an incomplete removal procedure (helm uninstall with the documented chart default cni.uninstall=false) leaves a CNI config that breaks every new pod sandbox cluster-wide until removed (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
  7. An engineering incident, not a security finding: the tested Linkerd configuration failed to carry the vLLM workload (every request returned 503); the root cause was not isolated, and the Linkerd version and proxy logs were not kept (L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)).

Background

Distributed LLM inference

vLLM serves an OpenAI-compatible API over HTTP on port 8000 by default, with streaming responses delivered as Server-Sent Events (SSE). Beyond completions it exposes tokenizer, embedding, scoring, and operational routes on the same port, plus a Prometheus /metrics endpoint. Authentication is optional and off by default; it is enabled with --api-key or VLLM_API_KEY, which requires a bearer token on the /v1 OpenAI routes. The check is not a blanket control for every route on the server. CVE-2026-48746 (GHSA-94f4-hr76-p5j6, CVSS 9.1, critical, published 2026-06-02) is an API-key bypass in the OpenAI API server: a crafted Host header makes AuthenticationMiddleware reconstruct a path that skips the /v1 API-key check. Affected versions are >=0.3.0 and <0.22.0; it is fixed in 0.22.0. Deployments behind an RFC-conforming reverse proxy such as nginx are not affected [19] [C].

When a model is split across GPUs on different nodes, vLLM can use Ray as its distributed executor. The Ray components relevant here are:

ComponentDefault portRole
GCS (Global Control Store)6379Cluster metadata, node and actor registry, gRPC
Dashboard and Job Submission API8265HTTP UI and REST API to submit arbitrary jobs
Ray Client server10001Remote driver connections
Raylet, object manager, worker RPC10002 and up, plus ephemeral portsTask scheduling, object transfer
Metrics exporter8080Prometheus

The Job Submission API accepts an entrypoint shell command and runs it on the cluster. Without authentication in front of it, any client that can reach 8265 can run code on Ray nodes. Pipeline parallelism (PP) places consecutive layer groups on different GPUs and passes activations between stages. Tensor parallelism across nodes all-reduces every layer and is far more latency sensitive. Disaggregated prefill and decode separates the two phases onto different workers and ships KV cache between them through a connector (in vLLM v0.11.0, P2pNcclConnector, built on NCCL plus a ZMQ side channel). NCCL has no built-in authentication or encryption over TCP sockets.

Kubernetes networking and NetworkPolicy

Kubernetes requires that every pod can reach every other pod without NAT. Isolation is opt-in through NetworkPolicy objects, which are enforced only if the CNI plugin implements them. On EKS, the VPC CNI enforces policy through the aws-eks-nodeagent (eBPF) when enableNetworkPolicy=true. A pod is isolated for ingress only once some policy selects it with policyTypes: Ingress, and for egress only once some policy selects it with policyTypes: Egress; the two directions are independent. NetworkPolicy operates on L3/L4 (IP, port, protocol, namespace and pod selectors); it provides no authentication of the caller and no encryption. A containerPort entry in a pod spec is informational metadata: omitting it does not stop a process from being reachable on that port.

The VPC CNI has a NETWORK_POLICY_ENFORCING_MODE setting. In the default standard mode, a new pod starts default-allow and policies are programmed in parallel with pod startup; in strict mode a new pod starts default-deny until its policies are programmed [18] [C]. This matters for any claim about a newly created pod, because in standard mode there can be a startup window during which a pod is reachable despite a matching deny policy.

Ray has optional token authentication from Ray 2.52.0 (RAY_AUTH_MODE=token). It is disabled by default and, when enabled, authenticates the dashboard, GCS and other control-plane services; KubeRay supports configuring it [4] [C]. Ray also supports TLS for its gRPC channels. Both are native alternatives or complements to network-layer controls.

Encryption in transit between pods can be added at the node level (WireGuard in Cilium or Calico, encrypting all node-to-node pod traffic) or at the workload level by a service mesh sidecar (Linkerd, Istio) that terminates mTLS and, for HTTP, proxies at L7.

Threat Model

Attacker. The attacker controls one ordinary pod in the same cluster, in a namespace unrelated to the inference stack. This models a compromised application container, a malicious dependency in a sidecar, or a tenant workload in a shared cluster. The pod runs as uid 1000, with every Linux capability dropped (CapBnd 0000000000000000) and no mounted ServiceAccount token (01-identity, probe-identity). The attacker has no Kubernetes API access, no node access, and no cloud credentials. The attacker can open TCP connections to pod IPs and Service names.

Assets.

AssetWhy it matters
Inference capacityFree GPU compute, and denial of service to legitimate callers
Prompt and completion contentMay contain regulated data; exposed via the API, via tokenizer routes, and on the wire
Workload telemetryRequest volume, token counts, and prefix-cache hit counters leak activity and may enable a cache side channel
Distributed control planeRay GCS and the Job API govern what code runs on GPU nodes
Inter-node tensorsActivations or KV cache derived from prompts, carried between GPUs

Out of scope. Container escape and node compromise; attacks on the Kubernetes API server, etcd, or cloud IAM; supply-chain attacks on images or model weights; model-level attacks (prompt injection, jailbreaks, extraction); side-channel exploitation (the prefix-cache channel is identified but not tested); attacks from outside the VPC; and any system not created for this study.

Method

Environment

ItemValueEvidence
ClusterDedicated EKS cluster sorami-lab, Kubernetes 1.35 (v1.35.8-eks), Amazon Linux 2023, containerd 2.2.704-gpu-nodes
NetworkDedicated VPC 10.99.0.0/16, region ap-southeast-2 (Sydney), no peering with other VPCsnetwork.tf
CNIAmazon VPC CNI with aws-eks-nodeagent --enable-network-policy=true. The addon exposes NETWORK_POLICY_ENFORCING_MODE, but its value was not recorded, so whether pods started default-allow (standard) or default-deny (strict) is unknown15-netpol-agent, 01-vpccni-encryption-options (variable name only)
CPU pool2 spot m5/m6i-class xlarge nodes (Ray CPU cluster, router, probe, scanners)00-nodes
Stage B GPU1x g6.xlarge spot, NVIDIA L4 24 GB00-target
Stage C/D GPUg5.xlarge on-demand (A10G 24 GB, PP stage 0, Ray head) plus g4dn.xlarge on-demand (T4 16 GB, PP stage 1, Ray worker)04-gpu-nodes
L2b session GPU2 x g5.xlarge on-demand (A10G + A10G), head and worker pinned to the same two nodes for all three runs02-decision, 00-target.log in each bench directory

Spot GPU capacity in Sydney was scarce during the study: the spot placement score was 1 of 10 for every g4dn, g5, and g6 size in every AZ, and a second spot GPU node never launched (spot-placement-scores). Stage C/D therefore ran on two on-demand nodes from a fallback node group, which is why the pair is heterogeneous.

Software under test

ComponentVersionConfiguration
KubeRay operator and ray-cluster chart1.7.1chart defaults (Stage A)
Ray2.52.0reported by /api/version. Token authentication (RAY_AUTH_MODE=token) is available in this version and was left at its default (off); Ray TLS was also off. Neither was evaluated
vLLM production-stack chart (router)0.1.12chart defaults
vLLM engine imagevllm/vllm-openai:v0.11.0 (tag only, digest not recorded; inside the CVE-2026-48746 affected range and exposed directly on its Service with no reverse proxy)Stage B: Qwen/Qwen2.5-7B-Instruct, bf16, --max-model-len 8192, --gpu-memory-utilization 0.90. Stage C/D: Qwen/Qwen2.5-1.5B-Instruct, --dtype half, --pipeline-parallel-size 2, --distributed-executor-backend ray, --enforce-eager, --gpu-memory-utilization 0.6, VLLM_ATTENTION_BACKEND=TORCH_SDPA, VLLM_USE_FLASHINFER_SAMPLER=0
ScannersTrivy 0.67.0, Kubescape v3.0.40, kube-linter 0.8.3, Checkov (version not logged)in-cluster Jobs on the rendered manifests

The Stage C/D engine settings are workarounds for the T4 (compute capability 7.5): the FlashInfer sampler crashed on its architecture check and the xformers backend raised NotImplementedError under PP; both are documented in stageC-distributed. The 1.5B model and fp16 were forced by the 16 GB T4 (no bf16 support).

Topology selection

The preferred Stage C topology was disaggregated prefill and decode with KV-cache transfer through P2pNcclConnector. It failed on both GPUs: the connector's pinned staging pool (tensor_memory_pool._allocate_pinned_memory) raised CUDA error: out of memory at kv_buffer_size values from 2e9 down to 1e8 and GPU memory utilisation from 0.85 down to 0.4 (06-p2p-oom-evidence, 07-topology-decision). We fell back to Ray-backed PP=2 across the two nodes, one GPU each. Tensor parallelism across nodes was not attempted because each node has one GPU and cross-node TP over TCP without EFA would measure the NIC rather than the security layers. Pod placement was pinned by kubernetes.io/hostname and was identical for every Stage D layer (same two node names in every 00-target.log).

Benchmark harness

The client is bench/bench.py in the test-harness/ folder of this repository. It is stdlib-only Python so the same bytes run in python:3.12-slim for every layer, delivered through a ConfigMap. It runs as a Kubernetes Job in the neighbour namespace, so the measured path is the same cross-namespace pod-to-Service path the attacker uses.

ParameterValue
Endpointstreaming POST /v1/completions on the engine Service, port 8000
Promptfixed English sentence repeated 8 times (about 89 prompt tokens)
Outputmax_tokens=128, ignore_eos=true, temperature=0
Token accountingserver-reported usage.completion_tokens via stream_options.include_usage
Warmup8 requests at concurrency 1, discarded
Sweepconcurrency 1, 4, 16, 32; 128 requests per level; closed loop (each worker thread issues its next request when the previous one finishes)
Timeout300 s per request

Metrics per request: TTFT (time to the first SSE chunk carrying text), E2E (request start to stream end), and TPOT = (E2E minus TTFT) / (tokens minus 1). Per level the harness reports p50, p95, p99, and mean of each, plus output tokens per second (sum of completion tokens divided by wall time) and requests per second. Percentiles are nearest-rank over 128 samples, so p99 is effectively the second-largest value and is sensitive to one or two outliers.

Raw per-request samples were never persisted. Inside the pod, bench/bench.py line 119 writes both the summary and the per-request list (json.dump({"summary": summary, "raw": results}, f)), but the Job command in manifests/stageB-bench-job.yaml line 49 prints only the ['summary'] object to the pod log, and the runner scripts split result-c*.json out of those RESULT log lines. The pod filesystem was discarded with the Job. Every result-c*.json in this study is therefore a 4-percentile-plus-mean summary. No latency distributions, confidence intervals, bootstrap statistics, or outlier analysis can be computed from the archived data, and none are reported. Both paths are in the local harness repo, which has no git remote and no commits, so they are cited as local paths only.

GPU utilisation was sampled with nvidia-smi inside the engine pods every ~20 s during each run. /metrics was scraped before, during, and after.

Hardening layers

One variable was changed per layer, cumulatively where stated.

LayerChangeArtifact
L0None (distributed baseline)manifests/stageC-ray-pp.yaml
L1Default-deny ingress in the engine namespace (ingress only; egress left open); allow engine pods to each other on all ports; allow port 8000 only from the neighbour and engine namespacesmanifests/stageD-L1-networkpolicy.yaml
L2Linkerd mTLS sidecars on both engine pods and the bench client; Ray ports 6379 and 10002 to 10010 skipped by the proxyLinkerd edge control plane, MESH_BENCH=1 in scripts/stageC-bench-run.sh
L2bCilium 1.20.2 chained on aws-cni with WireGuard pod-to-pod encryption; L1 removed and no API key for this run; different GPU pair, own L0 bracketmanifests/stageD-L2b-cilium-values.yaml, manifests/stageD-L2b-ray-pp.yaml (test-harness)
L3L1 plus engine --api-key from a Secret; client sends the bearer tokenVLLM_API_KEY secretKeyRef in manifests/stageC-ray-pp.yaml, BENCH_AUTH=1

The L1 allow of all ports between engine pods is deliberate: Ray and NCCL use ephemeral ports chosen at runtime, so a port list would break the pipeline. It also means L1 does not restrict traffic inside the trust boundary of the engine itself.

L1 is ingress-only. Every policy in manifests/stageD-L1-networkpolicy.yaml sets policyTypes: ["Ingress"] (the default-deny policy at line 22), and the comment at line 11 records that egress was deliberately left open so the engine could pull the model and resolve DNS. L1 therefore tests who can reach the engine. It does not show containment of a compromised engine, which would need egress policy and an egress test; neither was done.

Baseline provenance (important)

The L0 control for Stage D is not a fresh run. The four result-c*.json files and 10-table.md in stageD-00-L0 are byte-identical to those in stageC-06-bench. We verified this with SHA-256:

FileSHA-256 (both directories)
result-c1.json2157d70330b94ee1c7975c452eaddcc8f676fd7905609c96c1edb21467ccfedf
result-c4.jsoned3771e0a6f1b7933c870e61be62a89dc9501f9f820dd83f705b0bb15b4120cb
result-c16.json99fba85e51c6ce57ca05399aa2fbbc1802bd62939af0303b1d79be239e0b9ee7
result-c32.jsonc9dcdda0e8d751159af853b1692c19b9bc34f740fbf58535a44e3eea2b08fa09
10-table.mdac07e74eacda7cbae3206f25634ff2b29931ddab9d662c711997128b35fafeb9

The L0 directory also has none of the run artefacts every other layer has (00-target.log, 03-apply-job.log, 04-nvidia-smi-load.csv.log, 08-job-logs.log), and its files carry a local modification time of 2026-09-28T09:06:22Z, about 7.5 minutes after the Stage C job finished writing results at 08:58:55Z. The brief asked for L0 to be re-confirmed on the exact cluster state used for L1 onward; that re-run did not happen. The single L0 measurement was taken 2026-09-28T08:46:33Z to 08:58:55Z on the same two nodes and the same head and worker pods that L1 later used (L1 ran on those pods after one restart caused by applying the policy; L3 ran on fresh pods on the same nodes).

Consequence: every L0 to L1 and L0 to L3 delta in Results: hardening compares runs from different sessions (different pod lifetimes, and for L3 different pod instances). With n=1 per level, the comparison cannot separate a layer effect from session-to-session variance.

Run timeline (UTC, 2026-09-28)

WindowActivityEvidence
04:22 to 04:38Stage A inventory, neighbour scan, scanners (CPU only)stageA-06-neighbor, stageA-08-scanners
06:04:35 to 06:27:07Stage B benchmark (L4, 7B)stageB-07-bench
06:27Stage B neighbour re-probestageB-08-neighbor-reprobe
08:46:33 to 08:58:55Stage C distributed baseline (the only L0 run)stageC-06-bench
about 09:00 to 09:02Stage C cross-namespace re-inventory, Job API attempt, plaintext checkstageC-07-reinventory
09:21:46 to 09:33:46L1 benchmark; L1 re-probe log written 09:33:48stageD-01-L1
09:43:18 to 09:44:19L2 Linkerd benchmark attempt (all requests failed)stageD-02-L2
10:00:58 to 10:13:40L3 benchmark; auth proof log written 10:13:41stageD-03-L3
about 10:27GPU on-demand group scaled 2 to 0stageD-99-gpu-teardown
12:01GPU on-demand group scaled 0 to 2; 2 x g5.xlarge returnedstageD-04-L2-wireguard
12:13:11 to 12:21:07L2b session: L0-beforebench-L0-before
12:27:23 to 12:35:42L2b session: Cilium chaining plus WireGuardbench-L2b
about 12:41 to 13:09Stale 05-cilium.conflist breaks new pod sandboxes22-stale-cni-conflist-breakage
13:08:41Cilium reinstalled with encryption disabled and cni.uninstall=true23-cilium-reinstall-noenc
13:09:03Stale conflist renamed on every node (same study team, parallel control session)22-remove-stale-cilium-conflist
13:13:26Bracket-state check: Cilium agent running, Encryption: Disabled, no cilium_wg0, head pod not a Cilium endpoint24-bracket-state
13:13:40 to 13:22:00L2b session: L0-after (Cilium agent still installed for the whole window)bench-L0-after
13:22:22Final Cilium uninstall, cni-uninstall=true30-cilium-uninstall-final
13:24:54Zero GPU instances25-gpu-zero-verify

Timestamps without seconds come from local file modification times converted to UTC, not from in-log # UTC stamps.

Safe-probing rules

  • Probes targeted only pod IPs and Service names inside sorami-lab-* namespaces, taken from an explicit allowlist built from the Kubernetes endpoints list (allowlist.txt). No CIDR sweeps.
  • Stage A used read-only handshakes only: HTTP GETs, TCP connects, and TLS client hellos. No POST to the Ray Job API.
  • Stage C allowed exactly one synthetic Ray Job submission whose entrypoint printed a marker string and the hostname. No file writes, shell escapes, outbound network, or credential access.
  • vLLM state-changing routes (/scale_elastic_ep, response cancel, LoRA load/unload, sleep/wake) were enumerated from /openapi.json and not called.
  • KV and tensor channels: reachability and protocol identification only (connect, banner, TLS check). No injection, tampering, or decoding of model data.

Results: attack surface

Four default scanner configurations and the runtime surface

We rendered the vendor-default charts (KubeRay operator 1.7.1, ray-cluster 1.7.1, vLLM production-stack 0.1.12) into one 934-line manifest with 21 objects, including 1 RayCluster custom resource and 2 Deployments, and ran four scanners on it as in-cluster Jobs with their default settings. This section is about those four default configurations only. We did not write or test custom rules, admission-time policy engines that render custom resources, or other scanners. The properties in question (reachability from another namespace, raylet ports chosen at runtime, cleartext gRPC) are runtime and network properties, not necessarily violations of a static manifest rule, so a scanner that did not report them was not necessarily failing at its documented job.

ScannerFindingsMain themesRayCluster mentionsEvidence
Trivy 0.67.031 of 143 tests failed (3 critical, 7 high, 10 medium, 11 low)securityContext on router and operator, image tags and registry, operator ClusterRole on secrets, roles, services, networkpolicies0out-trivy
Checkov20 of 247 checks failedsecurityContext, image digest and tag, SA token automount, CKV2_K8S_6 "pods which lack an associated NetworkPolicy" on the operator and router only0out-checkov
Kubescape v3.0.401 control (C-0013 non-root, 2 resources)non-root containers; router and operator listed as highest-stake workloads0out-kubescape
kube-linter 0.8.35 lint errorsCPU request, dangling Service, :latest, read-only root FS, runAsNonRoot0out-kubelinter

Fairness of comparing the finding counts with the RayCluster mentions (offline re-analysis of the stored scanner outputs and the rendered manifest, no new scans). What each scanner parsed, from its own output: Trivy's findings name the two Deployments, the KubeRay operator ClusterRole and two Roles; Checkov's 20 failed checks name the operator and router Deployments and the Pod objects it derives from them; Kubescape's one control covers 2 resources, and the two workloads it lists as highest-stake are the same two Deployments; kube-linter's 5 errors name 4 Deployment findings and one Service. Every finding therefore came from built-in workload, RBAC and Service kinds. None of the four outputs names the RayCluster custom resource, so we conclude the tools did not evaluate its embedded pod templates, not that they evaluated them and found nothing. Would any rule plausibly have applied to a RayCluster? The rendered RayCluster (69 lines in ray-cluster-1.7.1.yaml) sets no securityContext, references images by tag and not digest (rayproject/ray:2.52.0), sets imagePullPolicy: IfNotPresent and does not set automountServiceAccountToken. Rules of the kind that fired on the two Deployments (securityContext, non-root, read-only root filesystem, image digest, pull policy, service account token mounting, and Checkov's missing-NetworkPolicy graph check) would plausibly have fired on those pod templates if a tool had expanded them. That is an inference from rule wording and the manifest, not a measurement. No rule in the four outputs concerns reachability from another namespace, runtime listening ports, or cleartext gRPC, and the RayCluster declares no ports at all, so we would not expect a manifest rule to report the exposure in this section. The comparison is therefore not like for like: the finding counts in the table are hygiene findings on the objects the tools parsed, and zero counts mentions of an object they did not parse. It supports "the RayCluster was not evaluated", and it does not support a general claim about what scanners can detect. Whether a scanner with custom rules, or one that renders custom resources, would report the exposure was not tested.

The zero-mention result is a grep -aicE 'raycluster|ray-head|ray-worker|headGroupSpec|workerGroupSpecs' over each scanner output (crd-blindspot); we extended the same grep to the kube-linter output and it also returned 0. The scanners evaluated the two Deployments and the RBAC objects. The RayCluster, which is what creates the head and worker pods and the Ray listeners, was not evaluated. Checkov's missing-NetworkPolicy check fired for the operator and router but not for the Ray pods, because those pods exist only after the operator reconciles the custom resource.

A second, separate observation is visibility drift between declared port metadata and runtime listeners. The Stage A pod specs declare only containerPort 8080 (metrics) on the Ray head and worker (container-ports). A full TCP connect scan from the neighbour found far more (nmap-full-tcp): 15 of 17 observed Ray listening sockets were absent from the pod's declared port metadata. containerPort is metadata, not a security boundary, and Kubernetes does not block undeclared ports, so this is not a bypass of any Kubernetes control. The finding is narrower: a review or tooling pipeline that treats declared ports as an inventory of runtime exposure has incomplete visibility, seeing here a small fraction of the listeners that were reachable.

PodDeclared portsOpen ports seen by neighbourAbsent from declared metadata
Ray head80806379, 8080, 8265, 10001, 36087, 36349, 44217, 44227, 52365, 57914, 62226 (11)10
Ray worker80808080, 36027, 36999, 45109, 52365, 57992 (6)5
KubeRay operator80808080, 80821
vLLM router8000, 900080000

A review that reasons from declared ports (manual review or port-based policy generators) sees 1 of 11 head ports. Stage C showed the same pattern on the GPU pipeline: the head listened on 23 TCP sockets and the worker on 15, while the manifest declared 4 head ports and none on the worker (02-head-listen-sockets, stageC-ray-pp.yaml).

No chart shipped any NetworkPolicy (netpol), all four Stage A pods automounted a ServiceAccount token, and none set runAsNonRoot (sa-automount). The KubeRay operator ClusterRole grants create, delete, get, list, update, and watch on secrets cluster-wide (kuberay-clusterrole-secrets); Trivy did flag this one (AVD-KSV-0041, critical).

Bar chart of findings per scanner on the Deployments and RBAC objects they parsed, with no scanner output naming the RayCluster resource

Figure 1: The default scanners did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis. Bars show what each tool reported on the objects it did parse.

Exposed unauthenticated endpoints

All results below are from the unprivileged neighbour pod.

StageSurfaceResultAuthTLSEvidence
A (CPU Ray)Ray dashboard 8265 /api/version200, Ray 2.52.0, commit, session namenonenohandshakes
ARay 8265 /api/jobs/200, job list (empty)nonenosame
ARay 8265 /nodes?view=summary200, 17,097 bytes: hostnames, IPs, CPU and memory per nodenonenosame
ARay GCS 6379, client 10001, ephemeral RPCTCP accepted, non-TLS replynot testednosame
ARay metrics 8080 head / worker200, 143,182 / 42,392 bytesnonenosame
AvLLM router /health, /v1/models, /metrics200nonenostageA-baseline
B (L4 engine)/v1/models, /v1/chat/completions, /v1/completions200, model listed, completions returnednoneno05-models, 06-chat, 07-completions
B/openapi.json26 routes on port 8000, incl. /tokenize, /detokenize, /v1/embeddings, /v1/responses/{id}, /scale_elastic_ep (listed, not each called)none configuredno13b-openapi-routes-parsed
B/tokenize, /detokenizetext to 16 token ids; ids [785,6722,315,9625,374] back to "The capital of France is"noneno08-tokenize, 09-detokenize
B/metrics66 metric families incl. prefix_cache_hits_total 41,552 of prefix_cache_queries_total 46,357noneno11-metrics-counters, 12-metrics-names
C (PP=2)vLLM GET /v1/models, POST /v1/completions200, 200noneno04-vllm-unauth
CRay GCS 6379, raylet 10002 to 10006TCP accepted, cleartext gRPC prefacenone observedno01-portscan, 03-plaintext
CRay Job API 8265connection refusedn/an/a02-rayjob-8265

The tokenizer routes turn any leaked token-id stream (logs, traces, or a captured tensor channel) back into text without credentials. The prefix-cache counters, combined with the ability to send prompts, suggest a cross-tenant probe of whether a guessed prefix was recently served. That side channel is inferred and was not tested.

The distributed control plane: cleartext and reachable

Wire evidence. From the neighbour, a raw TCP connect to Ray GCS 6379 on the head and to raylet 10002 on both head and worker returned 46 bytes immediately, beginning 00 00 18 04 00 00 00 00 00: an HTTP/2 frame header of length 0x18 (24 bytes), type 0x04 (SETTINGS), on stream 0. That is the server side of a cleartext HTTP/2 (h2c) connection, which is how gRPC runs without TLS. A TLS ClientHello to the same ports failed with SSL: WRONG_VERSION_NUMBER (03-plaintext, 01-plaintext-connect). A Redis PING to 6379 got the same SETTINGS frame rather than +PONG, confirming that modern Ray GCS is gRPC and not Redis (02-plaintext-gcs-ping).

Log excerpt: stageC-07-reinventory/03-plaintext.log (from the neighbour pod, byte samples shortened)

== plaintext peek ==
<head>:6379 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
<head>:10002 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
<worker>:10002 plaintext_recv=46B sample=b'\x00\x00\x18\x04\x00\x00\x00\x00\x00...'
== TLS handshake attempt (expect refused) ==
<head>:6379 TLS_REFUSED SSLError [SSL: WRONG_VERSION_NUMBER] wrong version number
<head>:10002 TLS_REFUSED SSLError [SSL: WRONG_VERSION_NUMBER] wrong version number

Observation

Both Ray listeners opened with a cleartext HTTP/2 SETTINGS frame and could not negotiate TLS. No gRPC message was decoded.

We did not capture inter-node pipeline traffic with tcpdump, and we did not decode any gRPC messages. The claim is therefore narrower than "activations are readable on the wire": the Ray control and coordination listeners accept unauthenticated cleartext gRPC connections from the neighbour pod, and a TLS client cannot even negotiate with them. The NCCL activation channel between PP stages was not separately identified or captured, so its encryption state is inferred from NCCL's documented defaults (TCP sockets, no TLS) and was not measured.

Job API and remote code execution: claim by claim.

ClaimDemonstrated?Basis
The Ray Job API can be exposed to another namespaceYesStage A: 8265 answered unauthenticated GETs from the neighbour pod (handshakes)
The default CPU Ray deployment exposed itYesStage A, KubeRay default image rayproject/ray:2.52.0: /api/jobs/ returned 200 with an empty list. No job was submitted there
The GPU vLLM deployment exposed itNoStage C: 8265 refused from the neighbour and from inside the head pod (02-rayjob-8265, 03-dashboard-http-inside)
Remote code execution against the GPU inference stackNoThe single permitted synthetic submission failed with connection refused. No code ran on any Ray node
Ray control ports were reachable cross-namespaceYesStage C: 6379 and 10002 to 10006 open on the head, 10002 on the worker, from the neighbour pod (01-portscan)

Job API result (negative for the GPU stack, image-specific). On the Stage A CPU RayCluster (KubeRay default image, ray[default]), 8265 answered unauthenticated GETs including /api/jobs/, so submission is very likely possible; per the safe-probing rules we did not POST there. On the Stage C GPU pipeline we made the one permitted synthetic submission: TCP to 8265 was refused and the submission failed with URLError [Errno 111] Connection refused (02-rayjob-8265). From inside the head pod, 127.0.0.1:8265 was also refused (03-dashboard-http-inside). The vllm/vllm-openai:v0.11.0 image ships minimal Ray, so the dashboard and job server never start. A cross-namespace Ray Client attempt on 10001 also failed (ray[client] absent; connection timeout) (02-crossns-ray-client-marker, 03-crossns-rayclient-10001).

Log excerpt: stageC-07-reinventory/02-rayjob-8265.log (the one permitted synthetic submission)

TCP <head>:8265 REFUSED (ConnectionRefusedError)
SUBMIT FAILED URLError <urlopen error [Errno 111] Connection refused>

So remote code execution via the Job API was not demonstrated on the GPU pipeline, and was not attempted on the CPU deployment. Whether code execution is possible depends on the image: any image with ray[default], which is what the KubeRay default image used in Stage A ships, exposes 8265. Whether a client with network access to GCS and raylet alone can schedule work was not tested.

Chart comparing declared container ports with observed listening sockets on the Ray head and worker pods

Figure 2: Declared versus observed listening ports. 10 of 11 head sockets and 5 of 6 worker sockets were undeclared.

Reachability matrix

Cross-namespace from the neighbour pod. "Open" means TCP connect succeeded.

Log excerpt: stageC-07-reinventory/01-portscan.log (head pod, L0)

=== target <head> ===
  6379 :: OPEN
  8000 :: OPEN
  8265 :: closed
  10001 :: closed
  10002 :: OPEN
  10003 :: OPEN
  10004 :: OPEN
  10005 :: OPEN
  10006 :: OPEN
Target (Stage C/D)PortPurposeL0 (Stage C)After L1After L3
Head (A10G)6379Ray GCSopenblockednot re-probed (L1 still in place)
Head8000vLLM APIopen, no authopen, no authopen; ordinary requests to /v1/models and /v1/completions get 401 without key; other routes and the CVE-2026-48746 bypass not tested
Head10002 to 10006raylet, object manageropenblockednot re-probed
Head8265Ray Job APIclosednot probed (denied by policy)not re-probed
Head10001Ray Clientclosednot probednot re-probed
Worker (T4)10002rayletopenblockednot re-probed
Worker6379, 8000, 10003 to 10006not listeningclosedblockednot re-probed

Sources: 01-portscan, 01-reprobe-after-L1, 01-auth-proof. The L1 re-probe log lists 6379, 8000, and 10002 to 10006 on both hosts; it does not include 8265, which was not listening at L0 anyway. Note that the L1 probe reports "BLOCKED" on the worker for ports that were already closed at L0; the probe cannot distinguish a dropped SYN from a port that is not listening, so only the ports that were open at L0 (6379 and 10002 to 10006 on the head, 10002 on the worker) are evidence that L1 works. All probes were steady-state probes of long-running pods; the startup window after a new pod is created (Kubernetes networking and NetworkPolicy) was not probed.

Grid showing Ray head and worker ports open from the neighbour pod at L0 and blocked after the L1 NetworkPolicy

Figure 3: Cross-namespace reachability of the Ray pods before and after the ingress NetworkPolicy.

Results: hardening

Single-node reference (Stage B, not comparable to Stage D)

For context only: one g6.xlarge (L4 24 GB), Qwen2.5-7B-Instruct bf16, vendor defaults, run 2026-09-28T06:04:35Z to 06:27:07Z. 128 of 128 requests succeeded at every level.

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.13617.469.2 / 71.8 / 78.857.3 / 57.3 / 57.47344.9 / 7352.3 / 7354.6
40.53167.9121.9 / 122.9 / 153.358.4 / 58.5 / 58.57538.4 / 7546.8 / 7567.2
162.07264.9135.3 / 152.4 / 175.559.8 / 59.9 / 59.97725.3 / 7744.8 / 7765.4
323.807487.2173.9 / 192.6 / 217.564.7 / 64.9 / 65.18393.2 / 8430.3 / 8451.5

Source: stageB-07-bench result-c*.json. The L4 ran at 99% median utilisation and a median 71.9 W against its 72 W cap across 63 load samples (04-nvidia-smi-load). Tail spread was tight (p99 within 1% of p50 for E2E at every level), which contrasts with the two-node pipeline below.

Distributed per-layer results

Same two nodes, same model and flags, same client. 128 requests per level. L0, L1, and L3 each had 0 failures across 512 requests. All numbers are from the per-level result-c*.json files.

L0: distributed baseline, no hardening (stageC-06-bench, reused as stageD-00-L0).

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.26934.438.2 / 47.7 / 51.428.9 / 31.3 / 32.73712.4 / 4026.4 / 4186.3
40.927118.759.7 / 386.4 / 7397.332.1 / 34.1 / 34.34138.9 / 4497.6 / 11748.2
163.198409.481.5 / 798.5 / 800.837.3 / 41.7 / 41.74807.4 / 5795.4 / 5797.1
323.353429.2474.6 / 10024.8 / 10028.854.0 / 56.7 / 131.17330.0 / 16881.9 / 16899.2

L1: NetworkPolicy (stageD-01-L1).

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.27134.638.0 / 47.9 / 49.428.7 / 31.0 / 31.83685.8 / 3979.5 / 4074.5
41.011129.560.4 / 91.9 / 405.430.7 / 32.4 / 32.83954.7 / 4176.5 / 4572.8
163.109398.092.7 / 986.9 / 988.237.2 / 42.5 / 42.54801.9 / 6079.6 / 6082.1
324.444568.8165.1 / 837.0 / 839.354.1 / 54.5 / 57.47031.6 / 7760.3 / 7776.2

L3: NetworkPolicy plus engine API key (stageD-03-L3).

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.26834.338.6 / 48.4 / 51.529.0 / 31.5 / 32.33717.9 / 4044.7 / 4138.7
40.952121.961.1 / 385.8 / 7405.530.8 / 32.5 / 32.73970.0 / 4487.8 / 11537.1
162.555327.073.4 / 10606.2 / 10624.337.0 / 40.3 / 41.84776.7 / 15584.5 / 15604.3
324.346556.2218.0 / 786.5 / 1028.254.0 / 61.5 / 61.57174.6 / 8597.5 / 8599.4

GPU utilisation under load (samples every ~20 s; min / median / max):

LayerSamplesHead A10G util %Worker T4 util %Head / worker peak memory MiB
L0340 / 12 / 150 / 37 / 6814,160 / 8,583
L1330 / 12 / 160 / 39 / 5914,160 / 8,557
L3350 / 12 / 160 / 37 / 6714,140 / 8,541

Sources: 04-nvidia-smi-load.csv.log in each layer directory. The T4 (PP stage 1) is consistently the busier device and is the likely pipeline bottleneck. Neither GPU approaches saturation, which fits a small model in eager mode where per-step scheduling and inter-stage transfer, not compute, dominate. GPU utilisation is indistinguishable across layers.

Performance comparison

Percent change versus L0, computed from the raw JSON. Positive on throughput means higher than L0; positive on latency means slower. These are single runs from different sessions (Baseline provenance (important)). No confidence intervals can be computed from n=1, and the numbers must be read as "observed difference", not "effect of the layer".

concLayerout tok/sTTFT p50TPOT p50E2E p50E2E p95
1L1+0.6%-0.5%-0.7%-0.7%-1.2%
1L3-0.3%+1.0%+0.3%+0.1%+0.5%
4L1+9.1%+1.2%-4.4%-4.5%-7.1%
4L3+2.7%+2.3%-4.0%-4.1%-0.2%
16L1-2.8%+13.7%-0.3%-0.1%+4.9%
16L3-20.1%-9.9%-0.8%-0.6%+168.9%
32L1+32.5%-65.2%+0.2%-4.1%-54.0%
32L3+29.6%-54.1%0.0%-2.1%-49.1%
anyL2 (Linkerd)failed: 0 of 128 ok at every leveln/an/an/an/a
1L2b (Cilium chaining + WireGuard), vs mean of same-session L0-3.5%+8.3%+3.6%+3.2%+3.1%
4L2b-3.4%+6.2%+3.5%+3.8%+4.0%
16L2b-4.5%+11.1%+5.3%+5.4%+3.4%
32L2b-6.5%+45.8%+8.4%+8.4%+4.3%

The L2b rows use a different baseline and different hardware (L2b: Cilium chaining plus WireGuard combined (single bracketed run)): they compare against the mean of two same-session L0 runs on 2 x A10G and are not comparable in absolute terms to the L1 and L3 rows. The L2b TTFT p50 deltas are small absolute changes (for example 66.7 to 97.2 ms at c=32) and the L0-to-L0 TTFT drift at c=32 was -24.1%, so they are not read as a difference.

How to read this table (L1 and L3 rows).

  1. Median per-token and median full-request (E2E) latency are stable. TPOT p50 is within 4.4% of L0 at every level for both layers, and E2E p50 is within 4.5%. This is the most stable signal in the data, because medians over 128 samples are insensitive to a handful of stalled requests.
  2. Throughput and tail latency swing in both directions by large amounts. At c=32 both hardened layers show about 30% more throughput than L0; at c=16 L3 shows 20% less. Neither NetworkPolicy (per-packet eBPF lookup) nor a bearer-token string comparison is a plausible cause of a 30% gain or a 20% loss when TPOT p50 moved by under 1%.
  3. The swings trace to a small number of stalled requests. L0 at c=32 has TTFT p95 of 10,024.8 ms against a p50 of 474.6 ms, and L3 at c=16 has TTFT p95 of 10,606.2 ms against a p50 of 73.4 ms; the same pattern of multi-second TTFT outliers appears at c=4 in L0 (p99 7,397.3 ms) and L3 (p99 7,405.5 ms) but not in L1. Because the harness is closed-loop and the wall time at c=16 and c=32 is only 29 to 50 s, a few requests that wait roughly 7 to 10 s before their first token stretch the wall clock and drag throughput down for the whole level. We did not identify the cause of these stalls (candidates include scheduler admission under PP, Ray actor scheduling, and eager-mode first-step cost); the stall pattern is independent of which layer is applied.
  4. The L0 run at c=32 is the outlier in this study. It happens to be the reused baseline, so every c=32 comparison inherits its stall.

Conclusion for RQ3 (performance). Within cross-session variance, with n=1 run per level, we observed no measurable steady-state difference for L1 (NetworkPolicy) or L3 (NetworkPolicy plus API key). We do not claim that either layer improves throughput. For L2b (Cilium chaining plus WireGuard combined), measured against a same-session bracket, the observed difference in this single bracketed run was 3.4% to 6.5% lower throughput and 3.2% to 8.4% higher median E2E; the c=1 and c=4 part is within L0-to-L0 drift, and the c=16 and c=32 throughput difference is outside it. With n=1 per condition, no per-request data, no confidence intervals, heterogeneous GPUs across sessions and a closed-loop client, we do not present these figures as a measured cost of encryption. The data cannot rule out a difference smaller than the run-to-run spread, which on this topology is at least several percent on median E2E and tens of percent on throughput and tail latency.

Security effect of each layer

L1: NetworkPolicy. After applying the three policies, the same probe script from the same neighbour pod turned every port that was open at L0 into BLOCKED, while 8000 stayed open as intended (01-reprobe-after-L1).

Command: apply the L1 policies (Stage D)

kubectl apply -f manifests/stageD-L1-networkpolicy.yaml

Log excerpt: stageD-01-L1/01-reprobe-after-L1.log (head pod, same probe from the same neighbour pod)

  6379 :: BLOCKED
  8000 :: OPEN
  10002 :: BLOCKED
  10003 :: BLOCKED
  10004 :: BLOCKED
  10005 :: BLOCKED
  10006 :: BLOCKED
TargetPort(s)L0L1
Head6379openblocked
Head10002 to 10006openblocked
Head8000openopen (allowed)
Worker10002openblocked

Operational effect: applying default-deny to the running Ray cluster forced one restart of both engine pods; the L1 target log shows RESTARTS 1 (3m41s ago) on the head and 1 (3m42s ago) on the worker at 09:21:46Z (00-target). The cluster re-formed under the policy and the benchmark ran clean. The mechanism of the restart (for example conntrack state for established Ray connections being dropped when the eBPF policy attached) was not investigated.

Observation

What L1 does not do: it does not authenticate callers on 8000, it does not encrypt anything, and it allows all traffic between engine pods, so a compromise of either engine pod still reaches the whole Ray control plane. It is ingress-only (policyTypes: ["Ingress"], line 22 of manifests/stageD-L1-networkpolicy.yaml; egress left open by design per the comment at line 11), so it does not show containment of a compromised engine: outbound connections from the engine to other namespaces, the instance metadata service, or the internet were neither restricted nor tested. The evidence is steady-state only; the VPC CNI enforcing mode was not recorded (Limitations), so a possible default-allow window for newly started pods is unmeasured.

L3: API key on top of L1. The engine was restarted with --api-key sourced from a Kubernetes Secret (01-auth-proof):

Log excerpt: stageD-03-L3/01-auth-proof.log

--- no auth ---
no-key /v1/models -> 401
no-key /v1/completions -> 401
--- wrong key ---
bad-key /v1/models -> 401
--- correct key ---
good-key /v1/completions -> 000
command terminated with exit code 28
good-key /v1/models -> 200
good-key /v1/completions -> 200
Call from neighbourBefore (Stage C)After L3
No key, GET /v1/models200401
No key, POST /v1/completions200401
Wrong key, GET /v1/models200401
Correct key, GET /v1/modelsn/a200
Correct key, POST /v1/completionsn/a200 (a first attempt timed out with curl exit 28, the retry returned 200)

Observation

What this shows, and what it does not. L3 is the weakest security result in this study. It shows that the normal key path returned 401 on two /v1 routes (/v1/models, /v1/completions) with no key or a wrong key, and 200 with the key. It does not show that API-key authentication protects vLLM, and it is not evidence of authentication closure, for three reasons:

  1. Route coverage. We did not test the other 24 routes. In vLLM the check applies to the /v1 routes, and /metrics, /health, /tokenize, and /detokenize may remain open. That is unverified here.
  2. Known bypass. The tested engine was vllm/vllm-openai:v0.11.0, which is inside the affected range (>=0.3.0, <0.22.0) of CVE-2026-48746 / GHSA-94f4-hr76-p5j6 (CVSS 9.1, published 2026-06-02): a crafted Host header makes AuthenticationMiddleware reconstruct a path that skips the /v1 API-key check [19] [C]. The advisory says deployments behind an RFC-conforming proxy such as nginx are not affected; ours was exposed directly on its Service with no proxy. The advisory was public on the experiment date (2026-09-28), so the tested version was publicly known to be bypassable at experiment time. We did not test the bypass.
  3. Cleartext key. The key travels in cleartext over plain HTTP, so any party that can observe pod traffic can capture it.

Any reuse of this configuration should run vLLM 0.22.0 or later, or put an RFC-conforming proxy in front, and then test every route with no key, a wrong key, and a correct key.

L2: encryption in transit

L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)

We treat this section as an engineering incident report, not a security finding: it says nothing about the attack surface in Results: attack surface. The intent was to encrypt and mutually authenticate the client to engine path with a service mesh, leaving everything else unchanged. Linkerd (edge channel; exact version not kept, and proxy logs not kept) installed cleanly: the destination, identity, and proxy-injector pods were Running, and both engine pods were meshed with an issued workload identity default.sorami-lab-stack-vllm.serviceaccount.identity.linkerd.cluster.local. Ray ports 6379 and 10002 to 10010 were excluded from the proxy because routing them through it deadlocked raylet on this build. With those exclusions ray status showed both nodes active, and a completion issued from inside the head pod worked (05-l2-outcome).

The benchmark client was then meshed too (linkerd.io/inject: enabled plus proxy-await) so that the client to 8000 hop would be mTLS end to end. Every request failed:

concok / requestswall serror
10 / 1280.36HTTPError 503: Service Unavailable
40 / 1280.13same
160 / 1280.13same
320 / 1280.13same

Source: stageD-02-L2 result-c*.json, 08-job-logs. The 0.13 s wall time for 128 requests means the client-side proxy rejected requests immediately; they never reached the engine. The operator notes record that the L7 HTTP route returned "route default.http: service unavailable", and that switching port 8000 to opaque (L4) mode produced connection resets before the inbound proxy on the head saw any traffic (05-l2-outcome). Those two observations come from the outcome note rather than from captured proxy logs, and the opaque-mode run did not produce a result file of its own.

Interpretation, with the limits stated. The tested Linkerd configuration failed to carry the vLLM workload; the root cause was not isolated. This is not evidence that Linkerd is incompatible with vLLM or with SSE in general: the Linkerd version and the proxy logs were not kept, Service versus pod-IP routing was not separated, and one-sided versus two-sided meshing was not compared, so the failure cannot be attributed to SSE, to vLLM, or to Linkerd as a technology. Plausible contributors include the interaction between the Service used by the client (the KubeRay-managed -openai Service) and Linkerd's endpoint discovery for pods whose Ray ports were skipped, protocol detection on port 8000, and the ordering of sidecar readiness against vLLM readiness. The finding we are confident in is narrow and operational: a team that plans to mesh a Ray-backed vLLM deployment should expect to debug it, and should not assume that "inject the proxy" preserves the serving path. Because the meshed path never carried a successful request, no mTLS performance difference is reported. We also do not report a literature figure in its place.

Security effect of L2a as deployed: none on the attack surface measured in Results: attack surface. The Ray control-plane ports were excluded from the proxy, so even a working L2a would have left GCS and raylet traffic in cleartext.

L2b: Cilium chaining plus WireGuard combined (single bracketed run)

Node-level WireGuard is designed to encrypt pod traffic that crosses nodes without touching the application or its ports. Here it was tested only as part of the combined Cilium chaining plus WireGuard configuration (L2b). Interface counters showed inter-node pod traffic carried by the WireGuard device. Ray GCS, raylet, object transfer and NCCL flows were not individually identified or captured, so that these flows are encrypted is an inference from the datapath, not a measurement. It was not installable in place during the main session, so it ran later the same day as a separate, scoped experiment (stageD-04-L2-wireguard, decision record 02-decision).

Setup. The VPC CNI addon v1.22.4-eksbuild.3 exposes no encryption setting (01-vpccni-encryption-options). We installed Cilium 1.20.2 in CNI chaining mode behind aws-cni (cni.chainingMode=aws-cni, native routing, no masquerade, kubeProxyReplacement=false, policyEnforcementMode=never, encryption.type=wireguard, nodeEncryption=false, enableRouteMTUForCNIChaining=true). VPC CNI kept IPAM and pod IPs; only the recreated engine pods became Cilium endpoints (10-cilium-install, 11-recreate-engine-under-cilium). The L1 policies were removed for all three runs and restored afterwards, and no API key was set, so encryption was the only intended variable.

Hardware change. On-demand capacity returned 2 x g5.xlarge (A10G + A10G) rather than the Stage C g5 + g4dn pair. Head and worker were pinned to the same two nodes (ip-10-99-17-131, ip-10-99-38-120) for all three runs. Removing the T4 bottleneck roughly tripled c=32 throughput (about 1180 against 429 tok/s), so absolute L2b numbers are not comparable to L0, L1, or L3 in Distributed per-layer results. Only the within-session deltas below are meaningful.

Order (UTC, 2026-09-28; one run per level, 128 requests per level, 128 of 128 ok at every level in all three runs):

RunBench windowState
L0-before12:13:11 to 12:21:07plain VPC CNI, no encryption
L2b12:27:23 to 12:35:42Cilium chained + WireGuard
L0-after13:13:40 to 13:22:00Cilium agent reinstalled and running with encryption disabled, no cilium_wg0; head pod not a Cilium endpoint; worker pod not checked (24-bracket-state). Not a proven Cilium-free baseline (see below)

L0-before (bench-L0-before):

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.39750.829.2 / 33.6 / 38.719.6 / 20.4 / 20.62522.0 / 2619.5 / 2655.0
41.458186.740.5 / 56.1 / 105.021.2 / 22.0 / 22.22730.1 / 2830.7 / 2860.3
165.313680.151.8 / 88.2 / 100.423.1 / 23.9 / 24.02984.8 / 3102.8 / 3104.9
329.2171179.875.8 / 144.2 / 149.526.6 / 27.9 / 27.93434.3 / 3615.4 / 3617.3

L2b, Cilium chaining plus WireGuard (bench-L2b):

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.38148.731.4 / 35.0 / 37.420.4 / 21.2 / 21.72618.4 / 2720.4 / 2783.1
41.400179.243.4 / 65.8 / 117.522.1 / 23.0 / 23.22853.9 / 2967.0 / 2991.4
164.986638.260.6 / 80.3 / 100.524.9 / 25.2 / 25.33216.2 / 3264.9 / 3270.0
328.7261116.997.2 / 125.8 / 158.128.4 / 28.8 / 28.93665.8 / 3759.8 / 3763.1

L0-after, drift bracket (bench-L0-after):

concreq/sout tok/sTTFT p50 / p95 / p99 msTPOT p50 / p95 / p99 msE2E p50 / p95 / p99 ms
10.39250.128.8 / 33.3 / 36.019.8 / 20.7 / 21.02550.2 / 2659.0 / 2696.5
41.441184.441.2 / 49.6 / 116.421.5 / 22.3 / 22.42766.7 / 2872.7 / 2880.1
165.131656.857.3 / 90.9 / 100.024.2 / 24.9 / 24.93119.8 / 3210.1 / 3214.2
329.4451209.057.5 / 150.7 / 178.725.8 / 27.1 / 27.23326.8 / 3596.7 / 3626.6

Difference: L2b against the mean of the two L0 runs. Drift: L0-after against L0-before.

concout tok/s L0 meanL2bdifferenceL0 driftE2E p50 L0 mean msL2b msdifferenceL0 drift
150.4548.7-3.5%-1.4%2536.12618.4+3.2%+1.1%
4185.55179.2-3.4%-1.2%2748.42853.9+3.8%+1.3%
16668.45638.2-4.5%-3.4%3052.33216.2+5.4%+4.5%
321194.41116.9-6.5%+2.5%3380.63665.8+8.4%-3.1%

TPOT p50 difference on the same basis is +3.6%, +3.5%, +5.3%, and +8.4%. The full metric-by-metric table, including the difference against L0-before alone, is in 50-wireguard-tax.

How to read this. The L0-to-L0 drift reaches about 3.4% on throughput (at c=16), so the c=1 and c=4 difference (-3.5% and -3.4%) is within drift, and the c=16 and c=32 difference (-4.5% and -6.5%) is outside it. On E2E p50 the drift is larger at c=16 (+4.5%), so only the c=32 latency difference (+8.4%) is clearly outside drift. The direction is consistent: L2b is slower than both L0 runs at every level on throughput, E2E p50, and TPOT p50. Tail percentiles show no clean signal (TTFT p95 at c=32 is lower under L2b than in either L0 run), so we claim no tail-latency difference. GPU medians were unchanged across the three runs (head 18%, worker 22 to 23%; 04-nvidia-smi-load.csv.log in each directory), so neither GPU was saturated and an inter-node network effect can surface in TPOT.

Confound. The L2b run had Cilium's eBPF datapath chained behind aws-cni plus WireGuard, so the measured difference is Cilium chaining plus WireGuard combined, not WireGuard alone. Nothing here isolates WireGuard. We tried a Cilium-without-encryption state, but helm upgrade --set encryption.enabled=false did not restart the agents and cilium-dbg still reported WireGuard active, so we discarded that state and did not benchmark it (17-cilium-disable-wg, 18-recreate-engine-cilium-noenc). The Cilium-only share of the difference was not isolated.

Threat to the L2b comparison: the closing bracket is not a proven Cilium-free baseline. 24-bracket-state (13:13:26Z) shows the Cilium 1.20.2 agent reinstalled and running during L0-after, with CNI Chaining: aws-cni, Encryption: Disabled, and no cilium_wg0 interface. Cilium was not removed until 30-cilium-uninstall-final at 13:22:22Z, after the L0-after sweep ended at 13:22:00Z. The head pod IP was absent from the Cilium endpoint list, which is consistent with the stale conflist having been renamed at 13:09Z so new pods used 10-aws.conflist, but the worker pod was not checked. So L0-after is a no-encryption control whose datapath state is only partly verified. If the worker ran as a Cilium endpoint, L0-after contains part of the Cilium datapath effect, and averaging it into the baseline would shrink the apparent difference. The L2b difference against L0-before alone is in 50-wireguard-tax and is larger at c=16 (-6.2% throughput).

Encryption evidence. Agent status on both GPU nodes reported Encryption: Wireguard [NodeEncryption: Disabled, cilium_wg0 ... Port: 51871, Peers: 3] and CNI Chaining: aws-cni, with recent handshakes to every peer (12-cilium-encryption-status). We then sent 50 MiB of a repeated ASCII marker from the head pod to the worker pod (13-wg-counter-proof):

Head-node counterBeforeAfterDelta
Bytes sent by head pod, received by worker podn/a55,189,000 / 55,189,00055.19 MB
cilium_wg0 tx_bytes651,04856,438,648+55.79 MB
ens5 tx_bytes641,459,248697,520,489+56.06 MB

The payload went into the WireGuard device (about 0.6 MB above payload, consistent with TCP/IP headers). Over the L2b sweep the head cilium_wg0 tx grew 572.5 MB and the worker cilium_wg0 rx grew 546.7 MB, so the pipeline traffic also crossed the tunnel (14-wg-counters-before-bench, 15-wg-counters-after-bench). This is interface-counter evidence that traffic was routed through cilium_wg0. It is not a packet capture showing ciphertext on the ENI: the Cilium agent image has no tcpdump and no privileged node shell was available, so the absence of the cleartext marker or HTTP/2 SETTINGS preface on ens5 was not observed directly.

Security scope. WireGuard protects the node-to-node wire. It does not close the cross-namespace reachability from Reachability matrix, because a neighbour pod's connection to 6379 is decrypted before delivery. L2b complements L1 and does not replace it.

Operational finding: an incomplete Cilium removal procedure broke new pods. This is a gap in our removal procedure, not Cilium misbehaving. Cilium was installed with the chart default cni.uninstall=false, which is documented in the Cilium Helm reference as controlling whether the CNI configuration is removed on agent shutdown [20] [C]. helm uninstall cilium at 12:41:16Z removed the agents but left /etc/cni/net.d/05-cilium.conflist on all four nodes, including the two CPU nodes that never ran an engine pod (19-cilium-uninstall, 22-remove-stale-cilium-conflist). Because it sorts before 10-aws.conflist, the container runtime kept invoking cilium-cni, and every new pod sandbox failed with unable to connect to Cilium agent ... dial unix /var/run/cilium/cilium.sock: connect: no such file or directory (22-stale-cni-conflist-breakage). Existing pods kept running. The breakage lasted from about 12:41Z to 13:09Z, when the file was renamed on each node. A later install with cni.uninstall=true uninstalled cleanly at 13:22:22Z, and a sanity pod on a CPU node networked afterwards (30-cilium-uninstall-final, 31-cni-sanity-pod). The lesson is about security-control lifecycle testing: removal and rollback are part of the control. Anyone trialling Cilium chaining on EKS needs an explicit cleanup step: set cni.uninstall=true and confirm the agents restarted with it before uninstalling, or remove the conflist on every node afterwards and verify with a throwaway pod per node pool. Separately, flipping encryption.enabled with helm upgrade changed the ConfigMap but not the running agents; confirm state with cilium-dbg status after any toggle.

Teardown. The GPU on-demand group was scaled from 2 to 0 (23-tf-plan-gpu0, 24-tf-apply-gpu0); the instance count reached 0 at 13:24:54Z (25-gpu-zero-verify), and both GPU groups were confirmed at desired 0 at 13:28:54Z (41-gpu-zero-verify). The L0-after sweep finished at 13:22:00Z, before any GPU instance drained. The L1 policies were restored at 13:23:06Z (32-restore-L1-netpol). The conflist rename (13:09Z) and the GPU scale-down (13:22 to 13:24Z) were done by the same study team, in a parallel control session alongside the one that ran the benchmarks; both are recorded in 60-encryption-proof-and-outcome. Neither overlaps a benchmark window.

Bar chart: Linkerd served 0 of 128 requests at each concurrency level while Cilium chaining plus WireGuard served all 128

Figure 4: Successful requests per concurrency level for the two encryption attempts.

Chart comparing the Cilium chaining plus WireGuard throughput difference with the drift between two unhardened baseline runs

Figure 5: The Cilium chaining plus WireGuard throughput difference against the drift between the two same-session baseline runs.

Summary of layers

LayerClosesProofObserved difference (n=1)Leaves open
L1 NetworkPolicy (ingress only)Cross-namespace ingress to Ray GCS and raylet from the tested neighbour pod (steady state)7 ports open at L0 (6 head, 1 worker) to blockedno measurable steady-state difference within variance; one pod restart when applied liveUnauthenticated 8000; cleartext everywhere; lateral movement between engine pods; all egress (no containment of a compromised engine); pod startup window unmeasured
L2a Linkerd mTLSIntended: encryption and identity on 8000identity issued, pods meshednot measurable: the tested configuration failed to carry the workload (503), root cause not isolatedEverything, as deployed; Ray ports excluded
L2b Cilium chaining plus WireGuard combinedCleartext inter-node pod traffic (Ray and NCCL flows inferred, not separately captured)agent status, peers, and cilium_wg0 counters (+55.79 MB for a 55.19 MB transfer); no ciphertext pcap3.4% to 6.5% lower throughput, 3.2% to 8.4% higher median E2E (single bracketed run, n=1; c=1/c=4 within drift; closing bracket not proven Cilium-free)Cross-namespace reachability (decrypted before delivery); callers not authenticated; removal needs an explicit CNI cleanup step
L3 API key (plus L1)Ordinary unauthenticated requests on two /v1 routes (/v1/models, /v1/completions)200 to 401 without key, 200 with keyno measurable difference within varianceNot authentication closure: 24 other routes untested; v0.11.0 is in the CVE-2026-48746 bypass range (not tested); key sent in cleartext

Discussion

Practical recommendations

Ordered by closed exposure per unit of effort, based on what this study measured.

  1. Apply namespace default-deny ingress before the engine starts. L1 was the only layer that blocked cross-namespace ingress to the Ray control plane from the tested neighbour, it showed no measurable steady-state difference, and it is plain Kubernetes. Allow the engine pods to reach each other on all ports (Ray picks ports at runtime) and allow the serving port only from named caller namespaces. Apply it with the RayCluster, not onto a live one, to avoid the restart seen here. Add an egress policy too if containment of a compromised engine matters; L1 here was ingress-only and gives no containment. On EKS, consider NETWORK_POLICY_ENFORCING_MODE=strict so new pods start default-deny; the startup window in the default standard mode was not measured here.
  2. Turn on Ray's own controls where the version supports them. Ray 2.52.0 already had token authentication (RAY_AUTH_MODE=token, off by default) covering the dashboard and GCS, and Ray supports TLS for gRPC. We did not evaluate either, so their effect on serving and their coverage are open questions, but they address the control-plane exposure at the application layer rather than relying on network placement alone.
  3. Encrypt below the application. The multi-node surface is many ports, most absent from declared port metadata, speaking cleartext gRPC and NCCL. A sidecar mesh that skips those ports (as L2a had to) does not encrypt them. Node-level encryption (WireGuard in the CNI) is designed to cover them without per-port configuration; here the combined Cilium chaining plus WireGuard configuration routed inter-node pod traffic through the tunnel (Ray and NCCL flows were not separately captured) and carried SSE streaming where the tested L7 sidecar configuration failed, and in this single bracketed run the encrypted configuration showed about 3% to 8% lower throughput and higher median latency on this PP=2 topology (Cilium chaining plus WireGuard combined, n=1). On EKS that currently means chaining Cilium behind the VPC CNI, which needs an explicit CNI cleanup step on removal (L2b: Cilium chaining plus WireGuard combined (single bracketed run)).
  4. Choose the image deliberately. Whether the Ray Job API exists depends on whether the image ships ray[default]. The minimal vLLM image did not; the KubeRay default image did. Treat an exposed unauthenticated Ray Job API on 8265 as a potential arbitrary-code-execution interface, and put it behind authentication or disable the dashboard. For the CPU deployment this rests on the documented Job API semantics; no job was submitted there.
  5. Inventory from the running pod, not from the manifest. Declared port metadata covered 1 of 11 listening ports on the Stage A Ray head. Use a socket listing inside the pod or an allowlist-driven connect scan from a neighbour namespace.
  6. Require an API key on the engine, stored in a Secret, on a vLLM version that is not affected by CVE-2026-48746 (0.22.0 or later), or behind an RFC-conforming reverse proxy. In this study the normal key path returned 401 on two /v1 routes; that does not show the key protects vLLM, and the version tested was publicly known to be bypassable. Test every route (including /metrics, /tokenize, /detokenize) with no key, a wrong key, and the correct key before relying on it.
  7. Scope the operator. The KubeRay operator ClusterRole with cluster-wide secrets access is the most privileged identity in the stack. Namespace-scoped operator deployment, where supported, reduces that blast radius.

Defence in depth ordering

Each layer covers a different gap and none is sufficient alone.

Question an attacker asksLayer that answers itMeasured here
Can I reach it?NetworkPolicy (L1)yes, ingress from the tested neighbour, steady state
Can it reach out if compromised?Egress NetworkPolicynot tested; L1 was ingress-only
Will it serve me?Authentication (L3), Ray token authnot established: the normal key path returned 401 on two /v1 routes; other routes and the CVE-2026-48746 bypass untested; Ray token auth not evaluated
Can I read or alter it on the wire?Encryption (L2)the tested L2a configuration failed (engineering incident); the combined L2b configuration routed inter-node pod traffic through WireGuard (counter evidence, no pcap; NCCL not separately captured)
Can I run code through it?Image choice, Job API auth, RBACJob API not listening on the GPU image, so no code execution demonstrated there (The distributed control plane: cleartext and reachable); not tested elsewhere

L1 first, because it is cheap and it is the only tested layer that shrinks the reachable Ray surface without touching Ray. Encryption and Ray's own controls next, because L1 still allows the intended callers and anything that shares their namespace. An API key is a weaker, supporting layer: the normal key path returned 401, the version and route-coverage caveats above apply, and without encryption the key itself is sent in cleartext on the same network that L1 only partially fences.

Network-layer versus L7 encryption for streaming inference

The two L2 attempts split along the layer they operate at. The Linkerd sidecar sits in the HTTP path and has to understand, route, and hold open every SSE response; in the tested configuration it returned 503 for every request (root cause not isolated), and it had to skip the Ray ports to keep raylet alive. WireGuard in the CNI (tested only as the combined Cilium chaining plus WireGuard configuration) operates below TCP and does not parse the application protocol; it carried every streaming request, and interface counters showed inter-node pod traffic crossing the tunnel without per-port configuration (Ray and NCCL flows were not separately captured, so their encryption is inferred). The observed difference in this single bracketed run was about 3% to 8% on throughput and median latency, larger at higher concurrency where more activation bytes cross nodes per second. Three caveats bound this conclusion: the measured difference includes the Cilium chaining datapath, the closing L0 bracket was not proven Cilium-free, and with a 1.5B model the bytes moved between stages per decode step are small, so a larger model will likely pay more (inference, not measured).

What the four default scanner configurations did not report, and why

Three structural reasons are consistent with the outputs.

  • Custom resources were not parsed. The four outputs name only built-in workload, RBAC and Service kinds. A RayCluster is a template for pods that the operator creates later, and none of the four tools mentioned it, so its pod spec and security context were not evaluated in these configurations.
  • Runtime ports are not in the metadata. Ray and NCCL bind most listeners at runtime. A tool that reads only containerPort cannot know they exist. containerPort is metadata, not a security boundary, and Kubernetes does not block undeclared ports.
  • Network posture is reported per object, not per flow. Checkov's CKV2_K8S_6 notes that a pod lacks a NetworkPolicy, which is useful, but no scanner output reported that a Ray control-plane listener was reachable from another namespace. That is a runtime property, and a reachability test such as the neighbour probe measures it.

These are observations about four tools at the versions tested, in default configurations, on one manifest. Custom rules, other scanners, newer versions, and admission-time policy engines that render custom resources may behave differently; we did not test them.

Limitations

Internal validity.

  • Reused baseline. The Stage D L0 files are byte-identical copies of the Stage C benchmark (SHA-256 identical for all five files; no run logs in the L0 directory). L0 was not re-run in the same session as L1 or L3. All L1 and L3 deltas are cross-session. L2b is the exception: it has fresh L0 runs before and after in the same session.
  • L2b confound: Cilium plus WireGuard. The L2b run added Cilium's chained eBPF datapath and WireGuard at the same time. A Cilium-without-encryption state was attempted but not benchmarked, because helm upgrade did not restart the agents and WireGuard stayed active. The reported L2b difference is therefore Cilium chaining plus WireGuard combined; the WireGuard-only share was not isolated.
  • L2b drift. The L0-to-L0 drift within the L2b session was up to about 3.4% on throughput and 4.5% on E2E p50 (both at c=16). The c=1 and c=4 L2b difference is within that drift. With n=1 per run, even the c=16 and c=32 difference rests on one sample per level.
  • L2b encryption proof is counter-based. The evidence is Cilium agent status, WireGuard peers and handshakes, and cilium_wg0 byte counters matching a known transfer. No packet capture showed ciphertext on the ENI.
  • L2b L0-after is not a proven Cilium-free baseline. 24-bracket-state shows the Cilium agent reinstalled and running with encryption off (no cilium_wg0) during L0-after, and the final uninstall (30-cilium-uninstall-final) was at 13:22:22Z, after the sweep. The head pod was shown not to be a Cilium endpoint; the worker pod was not checked. If the worker was a Cilium endpoint, the bracket mean includes some Cilium datapath effect and understates the L2b difference. Two actions in the L2b session (the conflist rename at 13:09Z and the GPU scale-down from 13:22 to 13:24Z) were taken by the same study team in a parallel control session; neither overlaps a benchmark window.
  • No raw per-request samples. bench/bench.py line 119 writes {"summary", "raw"} inside the pod, but manifests/stageB-bench-job.yaml line 49 prints only ['summary'], and only the printed line was kept. No distributions, confidence intervals, bootstrap estimates, or outlier analysis are possible from the archive, and the cause of the multi-second stalls cannot be investigated after the fact.
  • Tag-pinned artifacts. Images were referenced by tag (for example vllm/vllm-openai:v0.11.0), and runtime image digests were not recorded. The Linkerd version and the Checkov version were not recorded.
  • One run per level. Each layer and concurrency level has exactly one 128-request run, including all three L2b-session runs. There are no repeated trials, no confidence intervals, and no significance tests. The observed spread between runs (for example, throughput at c=32 from 429.2 to 568.8 tok/s across layers that should not affect throughput) indicates that run-to-run variance is large on this topology.
  • Different pod instances. L1 ran on the L0 pods after one restart; L2a and L3 ran on newly created pods on the same nodes. Pod-level warm state (allocator state, Ray actor placement, page cache) differs between layers.
  • Unexplained stalls. Multi-second TTFT outliers appear in L0 and L3 but not L1. Their cause is unknown, and they dominate tail and throughput comparisons.
  • Percentile resolution. With 128 samples, p99 is set by one or two requests.
  • Closed-loop client. Throughput is coupled to latency; a single stalled request lowers throughput for the level. An open-loop (fixed arrival rate) harness would separate the two.
  • Startup window. A search of every stored log for NETWORK_POLICY_ENFORCING_MODE found the name only in an EKS add-on option listing, with no value. The logs show --enable-network-policy=true on the aws-eks-nodeagent container and nothing about the enforcing mode. Ray pods created while the policy already existed (for example after a restart) were never probed during startup. L1 restarted both engine pods, and the probe ran on the re-formed, long-running pods. Whether a newly created Ray pod is reachable from another namespace during its startup window in this cluster is an open question that this study cannot answer.
  • Proof coverage. The L1 probe cannot tell a dropped SYN from a non-listening port, so only ports open at L0 count as evidence. L1 was ingress-only and no egress test was run. All probes were steady-state; the VPC CNI NETWORK_POLICY_ENFORCING_MODE value was not recorded (only the variable name appears in 01-vpccni-encryption-options), so a default-allow startup window under standard mode is unmeasured. The L3 proof covers ordinary requests to /v1/models and /v1/completions only; the CVE-2026-48746 Host-header bypass, which applies to the tested version, was not tested.
  • Ray native security not evaluated. Ray 2.52.0 supports token authentication and Ray supports TLS; both were off. The study does not compare network-layer controls against these application-layer controls.

External validity.

  • Heterogeneous GPUs. PP stage 0 on an A10G and stage 1 on a T4 is not a production configuration. The T4 gates throughput, and the settings needed to make it work (--enforce-eager, TORCH_SDPA, no FlashInfer sampler) change the performance profile.
  • Hardware change for L2b. The L2b session ran on 2 x A10G, not A10G + T4, with the same engine flags. c=32 throughput was about 1180 tok/s there against 429 tok/s on the Stage C pair. Absolute numbers are not comparable across the two sessions; only within-session deltas are. The percent difference may differ on the T4 pair, where the pipeline was slower and the network a smaller share of step time.
  • Small model. Qwen2.5-1.5B in fp16 moves far fewer bytes per step between stages than a 70B-class model. An encryption layer's overhead is expected to scale with inter-node bytes, so the L2b difference here may understate the difference for larger models (inference, not measured).
  • Topology. PP=2 over TCP, no EFA, no tensor parallelism across nodes, no disaggregated KV transfer (that path failed to start). The KV-transfer channel, which carries prompt-derived state, was not measured.
  • One region, one day, on-demand fallback. All GPU runs were in ap-southeast-2 on 2026-09-28. Stage C/D ran on on-demand nodes because spot capacity was unavailable; results say nothing about spot interruption behaviour.
  • Versions. vLLM v0.11.0, Ray 2.52.0, KubeRay 1.7.1, Linkerd edge (version not recorded), VPC CNI network policy agent. Defaults change across releases; for example, the Job API exposure depends on image contents, and the vLLM API-key behaviour differs between versions inside and outside the CVE-2026-48746 range.
  • Short prompts. About 89 prompt tokens and 128 output tokens per request. Long-context or long-output workloads change the ratio of prefill to decode and of control traffic to data traffic.

Construct validity.

  • "Plaintext" is shown for the Ray GCS and raylet listeners by the cleartext HTTP/2 preface and a failed TLS handshake. Encryption of the NCCL activation stream was not measured.
  • "No measurable difference" means no difference larger than the variance we could observe with n=1. It is not evidence of zero effect.
  • "L3 authenticated the engine" would be the wrong construct. What was measured is that ordinary requests on two /v1 routes were rejected without the key.
  • "L1 isolated the engine" would also be the wrong construct. What was measured is steady-state ingress blocking from one neighbour pod to specific ports.

Citation confidence is marked. [C] means the source, its identifier and the detail cited were checked (all entries were re-checked on 2026-09-29).

Ray exposure. Ray's documentation states that Ray assumes a trusted network and that the dashboard and Job API must not be exposed to untrusted clients [1] [C]. CVE-2023-48022 describes unauthenticated remote code execution via the Ray Jobs API; Anyscale disputes it as a vulnerability on the grounds that Ray is designed to run inside a trusted boundary [2] [C]. Oligo Security reported in-the-wild exploitation of internet-exposed Ray clusters under the name ShadowRay in 2024 [3] [C]. Our contribution is the in-cluster version of the same exposure: not internet-facing, but reachable from other pods in a flat Kubernetes network unless a NetworkPolicy blocks it, and dependent on image contents. Ray token authentication exists from Ray 2.52.0 (RAY_AUTH_MODE=token), is disabled by default, and authenticates the dashboard, GCS and other control-plane services; KubeRay supports configuring it [4] [C]. Ray 2.52.0 was the version under test, and we ran it with token authentication off and did not evaluate it.

vLLM API authentication. CVE-2026-48746 / GHSA-94f4-hr76-p5j6 is a critical (CVSS 9.1) API-key bypass in the vLLM OpenAI API server: a crafted Host header makes AuthenticationMiddleware reconstruct a path that skips the /v1 check. It affects >=0.3.0, <0.22.0, is fixed in 0.22.0, was published 2026-06-02, and does not affect deployments behind an RFC-conforming proxy such as nginx [19] [C]. The tested v0.11.0 engine was exposed directly and is in range; we did not test the bypass.

Inference-engine advisories. vLLM has published advisories for unauthenticated deserialisation on its distributed communication paths, including the PyNcclPipe KV-transfer service (CVE-2025-47277) [5] [C] and ZeroMQ sockets in the multi-node V0 engine (CVE-2025-30165) [6] [C]. A 2025 report by Oligo Security, which named the pattern ShadowMQ, described the same unsafe ZeroMQ plus pickle pattern across several inference frameworks including vLLM and SGLang [7] [C]. These advisories concern code paths we did not exercise (the disaggregated connector failed to start, and we did not send data to any tensor channel). They support the general point that inter-node inference channels are an attack surface and not just a performance concern.

LLM serving systems and side channels. vLLM's PagedAttention design [8] [C] and Ray [9] [C] are the systems under test. Automatic prefix caching creates shared state across requests; timing side channels on shared KV and prefix caches in LLM serving have been studied [10] [C]. Our observation that /metrics exposes prefix-cache hit counters unauthenticated is a possible amplifier of such channels; we did not test it.

Kubernetes NetworkPolicy. The Kubernetes documentation defines NetworkPolicy semantics and states that enforcement depends on the network plugin [11] [C]. Budigiri et al. evaluated the performance and security of NetworkPolicy implementations across CNIs [12] [C]. Our L1 result, no measurable steady-state difference on the VPC CNI eBPF agent, is consistent in direction with the expectation that L3/L4 filtering is cheap, but n=1 does not allow a quantitative comparison.

Service mesh and encryption overhead. Zhu et al. dissected the latency and CPU overhead of sidecar service meshes and attributed much of it to L7 processing [13] [C]. Istio and Linkerd both publish performance measurements [14] [C]. WireGuard's design and performance are described by Donenfeld [15] [C], and Cilium documents transparent WireGuard encryption between nodes [16] [C]. We found no prior measurement of mesh or node encryption overhead specifically on Ray-backed multi-node LLM inference; L2b gives one bracketed data point for this testbed (about 3% to 8%, Cilium chaining plus WireGuard combined, n=1). The L2a failure is a configuration-specific negative result, not an overhead result and not a demonstrated incompatibility.

Static analysis of Kubernetes manifests. Trivy, Checkov, Kubescape, and kube-linter are widely used open-source configuration scanners [17] [C]. We are not aware of a published comparison focused on their handling of operator custom resources; that comparison would be useful and is not claimed here.

Ethics and Responsible Disclosure

  • All testing ran in a dedicated Amazon EKS cluster and a dedicated VPC (10.99.0.0/16) in AWS, created for this study and used for nothing else. Only our own pods were probed.
  • No third-party systems, no internet-facing services, and no other tenants were probed. Every probe target came from an allowlist of our own pod IPs; no CIDR ranges were scanned.
  • Exploitation was limited to one benign synthetic Ray Job submission (marker print and hostname), which failed because the endpoint was not present. No state-changing vLLM routes were called. No tensor or KV data was injected, tampered with, or decoded. Prompts contained only synthetic text (a fake record number in the tokenizer test).
  • The behaviours reported (Ray's trusted-network assumption, vLLM's optional authentication, charts shipping no NetworkPolicy) are documented vendor defaults rather than new vulnerabilities, so no new coordinated disclosure is required. The vLLM API-key bypass relevant to the tested version (CVE-2026-48746) was already public and fixed upstream. If the Linkerd failure is reproduced with a minimal, version-pinned configuration and a root cause, it will be reported to the relevant upstream issue tracker as a bug, not a security issue.
  • Logs were redacted of names, emails, account and principal IDs, public IP addresses, and credentials before publication. GPU capacity was scaled to zero when the work finished.

Future work

Follow-ups drawn from the two independent reviews (Appendix D and Appendix E), ordered by how much they would change the conclusions.

  1. Route-by-route authentication matrix across vLLM versions. Enumerate every route from /openapi.json and call each with no key, a wrong key, and the correct key, on v0.11.0, a patched release (0.22.0 or later), and each behind an RFC-conforming proxy. Reproduce the CVE-2026-48746 regression only in an isolated environment. Classify each route as inference-capable, administrative, or information-disclosing.
  2. Ray native security versus network controls. A factorial of Ray token authentication (RAY_AUTH_MODE=token) and Ray TLS, crossed with NetworkPolicy and WireGuard, measuring both security effect (rejection, handshake behaviour, reachability) and serving difference (throughput, TTFT, TPOT, E2E, CPU).
  3. VPC CNI startup-window measurement. Probe a newly created target pod continuously from an unrelated namespace while recording sandbox creation, PolicyEndpoint programming, and Ready state; repeat many times under NETWORK_POLICY_ENFORCING_MODE=standard and under strict, and report the exposure-window distribution. No GPU needed.
  4. A/B/C Cilium decomposition. A: plain VPC CNI. B: Cilium chained with encryption disabled (with agents restarted and the state confirmed on every engine pod). C: the same Cilium plus WireGuard. Randomise the order across at least 10 blocks on identical nodes so that B minus A estimates the datapath difference and C minus B estimates the incremental WireGuard difference, with confidence intervals over blocks.
  5. Egress tests. Add egress NetworkPolicy and test outbound reach from a compromised-engine vantage point (other namespaces, instance metadata, the internet, model stores).
  6. Digest-pinned artifact manifest. Record runtime image digests, chart digests, model and tokenizer revisions, scanner, Linkerd, Cilium, kernel, driver and CNI versions in one machine-readable manifest with SHA-256 sums, and keep the harness under version control.
  7. Persist raw samples. Change the bench Job so the full {"summary","raw"} file leaves the pod (print it, or write it to a volume), so every run keeps per-request TTFT, TPOT and E2E for distributions, bootstrap intervals, and stall analysis.
  8. Linkerd root cause. Repeat L2a with a pinned Linkerd version, kept proxy logs, and a controlled split of Service versus pod-IP targeting, protocol detection versus opaque ports, and one-sided versus two-sided meshing. This is an engineering question, not a security one.
  9. What the exposed Ray ports allow beyond reachability. Not tested: whether a client with network access to GCS and raylet alone can read cluster state, schedule work or execute code on the GPU inference deployment. This needs an approved, isolated experiment.
  10. Ray native authentication and NetworkPolicy together. Not tested: whether NetworkPolicy is still needed when Ray token authentication or TLS is configured, and what each closes on its own (see item 2).
  11. Scanner fairness. Not tested: the same manifests with custom rules, with admission-time policy engines that render custom resources, and with scanners that parse the RayCluster resource, so that a like-for-like comparison of runtime-exposure detection is possible. First step: render the RayCluster into its resulting head and worker pod specs and rerun the same four scanners on them. The comparison in Four default scanner configurations and the runtime surface is not like for like.
  12. Newly created Ray pods during startup. Not tested (see item 3): probe pods created after the policy exists, with the enforcing mode recorded first.
  13. Job API and code execution on the GPU stack. Not demonstrated. If the Job API is enabled in a GPU deployment, test it only in an isolated environment with the same safe-probing rules.

Conclusion

In the tested KubeRay and vLLM topology, distributed inference introduced Ray control-plane listeners that were not represented by the workload's declared port metadata, and an ordinary pod in an unrelated namespace could reach them: Ray GCS and raylet spoke cleartext gRPC on ports that are mostly absent from the pod's declared port metadata. This is one cluster, one topology and one version set, and the reach was shown from one constructed neighbour pod. An ingress-only default-deny NetworkPolicy blocked cross-namespace ingress from that neighbour to the Ray control-plane ports in steady state; it gives no egress containment, and the pod startup window was not tested. The Ray Job API was exposed on the default CPU Ray image and was not listening in the GPU vLLM deployment, so remote code execution against the GPU inference stack was not demonstrated. Four default scanner configurations did not inspect the RayCluster custom resource, so the generated Ray pods were not represented in their analysis and the runtime network exposure was not identified; custom rules were not tested, and these are runtime and network properties that a reachability test measures. The weaker results follow. An engine API key turned ordinary unauthenticated requests on two /v1 routes into 401, which shows the normal key path works and does not show that API-key authentication protects vLLM: the other routes were not tested and the tested v0.11.0 was, at experiment time, publicly known to be bypassable through CVE-2026-48746 (not tested here). In a single bracketed run, Cilium chaining plus WireGuard combined carried every streaming request and routed inter-node pod traffic through the tunnel, with about 3% to 8% lower throughput and higher median latency, the low-concurrency part inside session drift and the closing bracket not proven Cilium-free; WireGuard alone was not isolated and NCCL traffic was not separately captured. Ray's own token authentication and TLS, available in the tested version, were not evaluated and are the obvious next comparison. Two engineering incidents are recorded for completeness: the tested Linkerd configuration failed to carry the vLLM workload (root cause not isolated), and an incomplete Cilium removal procedure broke new pods on every node until a stale CNI config was cleaned up. The single most important methodological lesson is that performance comparisons on a two-GPU pipeline need same-session baselines, repeated trials, and persisted per-request samples; only the L2b session had the first, and none of the runs had the other two.

Sorami's view

This section is opinion. It is how we would prioritise the fixes for a team running distributed inference.

Treat the Ray network as internal plumbing that must never be reachable from other workloads. Apply an ingress default-deny NetworkPolicy to the Ray namespace first, because it closed the most for the least effort in our test. Then add node-level encryption and check Ray's own authentication. Upgrade vLLM past 0.22.0 before trusting its API key, which is a weak layer on its own.

The four default scanner configurations we ran did not report this surface, and reachability is a runtime property. A reachability test from a neighbour pod measures it. That is part of our AI production readiness review and our cloud penetration testing.

What to do now

  • Apply an ingress default-deny NetworkPolicy to every Ray namespace.
  • Allow only the router to reach the vLLM API port.
  • Evaluate Ray token authentication, which is off by default.
  • Probe your Ray ports from a pod in another namespace.
  • Encrypt inter-node pod traffic with WireGuard or IPsec.
  • Upgrade vLLM to 0.22.0 or later, then set an API key.

Our AI on Kubernetes Helm chart research covers the single-chart defaults behind this stack. The practical guide gives the steps in order, and the guides index has related material.

Frequently asked questions

Is Ray secure to run on a shared Kubernetes cluster?

Not with defaults. The Ray documentation says a Ray cluster should run only inside a trusted network. In our one-cluster test the Ray control ports accepted cleartext gRPC from a pod in an unrelated namespace. Ray token authentication exists from Ray 2.52.0 but is off by default and was not evaluated here.

Do Kubernetes security scanners detect exposed Ray ports?

Not in the default configurations we tested. Trivy, Checkov, Kubescape and kube-linter did not inspect the RayCluster custom resource, so the Ray pods it generates were not represented in their analysis. Reachability is a runtime property, and custom rules were not tested. 15 of 17 observed Ray listening sockets were not declared as container ports, which are metadata only, so a review that relies on them has incomplete visibility.

Does a vLLM API key protect the server?

Not on its own. With an API key set, ordinary requests to /v1/models and /v1/completions got 401 instead of 200, which shows the normal key path works and does not show the API is protected. Other routes were not tested. The tested version is in the affected range of CVE-2026-48746, a Host header bypass fixed in vLLM 0.22.0, so upgrade before relying on the key.

How much did Cilium with WireGuard change LLM inference speed?

In one same-session bracket, Cilium chaining plus WireGuard combined showed 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median full-request latency. The difference rose with concurrency. It does not isolate WireGuard, and each figure is a single run on a small 1.5B model.

What is the cheapest fix for an exposed Ray cluster?

An ingress default-deny NetworkPolicy. In our test it blocked every Ray port that was open to the neighbour pod in steady state. The startup window for newly created pods was not tested, and the performance comparison is a single run per level.

Appendix: References

All identifiers and URLs below were checked on 2026-09-29.

  1. Ray Project. "Security" and "Configuring and managing the Ray Dashboard" in the Ray documentation. [C]
  2. NIST National Vulnerability Database. CVE-2023-48022 (Ray Jobs API, disputed by vendor). https://nvd.nist.gov/vuln/detail/CVE-2023-48022 [C]
  3. A. Lumelsky et al., Oligo Security. "ShadowRay: First Known Attack Campaign Targeting AI Workloads Actively Exploited In The Wild", 26 March 2024. Verified 2026-09-29. [C]
  4. Ray Project. "Ray token authentication" (available from Ray 2.52.0, RAY_AUTH_MODE=token, disabled by default); and "KubeRay authentication", https://docs.ray.io/en/releases-2.52.0/cluster/kubernetes/user-guides/kuberay-auth.html. Verified 2026-09-29. [C]
  5. vLLM Project. GHSA-hjq4-87xh-g4fv / CVE-2025-47277, "Remote Code Execution via PyNcclPipe Communication Service" (V0 engine PyNcclPipe KV-transfer only, affected >=0.6.5 <0.8.5, fixed 0.8.5), published 20 May 2025. Verified 2026-09-29. [C]
  6. vLLM Project. GHSA-9pcc-gvx5-r5wm / CVE-2025-30165, "Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration" (ZeroMQ SUB socket plus pickle in the multi-node V0 engine; not fixed, V0 off by default since 0.8.0), published May 2025. Verified 2026-09-29. [C]
  7. A. Lumelsky, Oligo Security. "ShadowMQ: How Code Reuse Spread Critical Vulnerabilities Across the AI Ecosystem", November 2025 (unsafe ZeroMQ recv_pyobj/pickle pattern in Meta Llama Stack, NVIDIA TensorRT-LLM, vLLM, SGLang, Modular Max Server and others). Verified 2026-09-29. [C]
  8. W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, I. Stoica. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023. [C]
  9. P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, I. Stoica. "Ray: A Distributed Framework for Emerging AI Applications." OSDI 2018. [C]
  10. X. Zheng et al. "InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks". arXiv:2411.18191, November 2024; and L. Song, Z. Pang, W. Wang et al. "The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems." arXiv:2409.20002, September 2024, https://arxiv.org/abs/2409.20002 (published in IEEE Transactions on Information Forensics and Security, vol. 20, 2025). Verified 2026-09-29. [C]
  11. Kubernetes documentation. "Network Policies". [C]
  12. G. Budigiri, C. Baumann, J. T. Muehlberg, E. Truyen, W. Joosen. "Network Policies in Kubernetes: Performance Evaluation and Security Analysis". EuCNC/6G Summit 2021, pp. 407-412. Verified 2026-09-29. [C]
  13. X. Zhu, G. She, B. Xue, Y. Zhang, Y. Zhang, X. K. Zou, X. Duan, P. He, A. Krishnamurthy, M. Lentz, D. Zhuo, R. Mahajan. "Dissecting Overheads of Service Mesh Sidecars". ACM SoCC 2023, pp. 142-157. Verified 2026-09-29. [C]
  14. Istio documentation, "Performance and Scalability"; Linkerd, published benchmark posts. [C]
  15. J. A. Donenfeld. "WireGuard: Next Generation Kernel Network Tunnel." NDSS 2017. [C]
  16. Cilium documentation. "WireGuard Transparent Encryption". [C]
  17. Project pages: Trivy https://github.com/aquasecurity/trivy, Checkov https://github.com/bridgecrewio/checkov, Kubescape https://github.com/kubescape/kubescape, kube-linter https://github.com/stackrox/kube-linter [C]
  18. Amazon EKS User Guide. "Configure network policy" (NETWORK_POLICY_ENFORCING_MODE: default standard starts new pods default-allow until policies are programmed; strict starts default-deny). Verified 2026-09-29. [C]
  19. GitHub Advisory Database. GHSA-94f4-hr76-p5j6 / CVE-2026-48746, vLLM OpenAI API server authentication bypass via crafted Host header, CVSS 9.1, affected >=0.3.0 <0.22.0, fixed 0.22.0, published 2026-06-02. https://github.com/advisories/GHSA-94f4-hr76-p5j6; OSV https://osv.dev/vulnerability/CVE-2026-48746. Verified 2026-09-29. [C]
  20. Cilium documentation. "Helm Reference" (cni.uninstall, default false). The default is documented; the URL was not re-fetched on 2026-09-29. [C]

Appendix A: Reproduction

The harness is in the test-harness/ folder of this repository. Paths below are relative to it. Copy infra/terraform.tfvars.example to infra/terraform.tfvars and set operator_cidr to your own /32. Every step that mutates infrastructure must be scoped to a dedicated, isolated account or VPC; do not run the neighbour probe in a shared cluster.

  1. Infrastructure. infra/ (Terraform): EKS 1.35 with VPC CNI enableNetworkPolicy=true, a CPU spot group, a GPU spot group (tainted sorami-lab/gpu=true:NoSchedule), and a GPU on-demand fallback group (max 2 for Stage C). Scale GPU groups with -var gpu_desired=N -var gpu_ondemand_desired=M.
  2. GPU device plugin. NVIDIA device plugin Helm release restricted to the GPU pool.
  3. Stage A. Install KubeRay operator 1.7.1, ray-cluster 1.7.1, and vLLM production-stack 0.1.12 with default values. Render them to manifests/rendered-defaults/all-defaults.yaml and run manifests/scanner-jobs.yaml.
  4. Neighbour probe. manifests/neighbor-probe.yaml: namespace sorami-lab-neighbor, uid 1000, all capabilities dropped, no ServiceAccount token. Build the allowlist from the endpoints list, then run nmap -sT -Pn -p- --open against only those IPs and scripts/handshake.sh for read-only handshakes.
  5. Stage B. manifests/stageB-vllm-engine.yaml (Qwen2.5-7B-Instruct, vllm/vllm-openai:v0.11.0), then scripts/stageB-bench-run.sh and scripts/stageB-reprobe.sh.
  6. Stage C. manifests/stageC-ray-pp.yaml with node names pinned to your two GPU nodes. Benchmark: scripts/stageC-bench-run.sh <logdir> "<label>".
  7. Stage D. L1: kubectl apply -f manifests/stageD-L1-networkpolicy.yaml, then re-probe and benchmark. L2a: install Linkerd, annotate the engine pods (skip Ray ports 6379 and 10002 to 10010), run the bench with MESH_BENCH=1. L2b: remove the L1 policies, run a fresh L0, helm install cilium cilium/cilium --version 1.20.2 -n kube-system -f manifests/stageD-L2b-cilium-values.yaml (set cni.uninstall=true from the start), recreate the engine pods, confirm cilium-dbg status shows WireGuard, run the bench, then remove Cilium, confirm no 05-cilium.conflist remains on any node, recreate the engine pods, and run L0 again. L3: create Secret sorami-lab-vllm-apikey (key apikey), restart the RayCluster, run the bench with BENCH_AUTH=1.
  8. Method fixes this study did not have: re-run L0 immediately before each layer in the same session, repeat each level in at least 10 randomised blocks, alternate layer order (for example L0, L1, L0, L3, L0), persist the per-request raw list that bench/bench.py already writes inside the pod (the Job at manifests/stageB-bench-job.yaml line 49 currently prints only summary), and report medians with bootstrap confidence intervals over blocks. Record image digests, not only tags. Consider an open-loop arrival-rate client to decouple throughput from stalls.
  9. Teardown. Scale both GPU groups to 0 and delete the sorami-lab-* namespaces; terraform destroy removes the cluster and VPC.

Appendix B: Artifact list

Findings summaries: stageA-baseline, stageB-baseline, stageC-distributed, stageD-hardening-tax. Repository guide: README.

ClaimRaw artifact
Rendered defaults scanned, scanner outputsstageA-08-scanners
Zero RayCluster mentionscrd-blindspot.log
Declared portscontainer-ports.log
Open ports, Stage Anmap-full-tcp.log
Ray dashboard, GCS, metrics handshakeshandshakes.log
Probe identityprobe-identity.log, 01-identity.log
RBACstageA-07-rbac
Stage B benchmark JSONstageB-07-bench
Stage B routes, tokenizer, metrics, TLS, NetworkPolicy absencestageB-08-neighbor-reprobe
Spot capacity scarcityspot-placement-scores.log
Stage C nodes and instance types04-gpu-nodes.log
Disaggregated connector OOM06-p2p-oom-evidence.log, 07-topology-decision.md
Listening sockets in engine pods02-head-listen-sockets.log
Cross-namespace port scan, Stage C01-portscan.log
Job API submission attempt02-rayjob-8265.log, stageC-04-rayjob
Cleartext gRPC preface, TLS refused03-plaintext.log, stageC-05-plaintext
Unauthenticated vLLM, Stage C04-vllm-unauth.log
L0 benchmark (only run)stageC-06-bench
L0 copy used in Stage DstageD-00-L0
L1 benchmark and re-probestageD-01-L1
L2a Linkerd attemptstageD-02-L2
L3 benchmark and auth proofstageD-03-L3
L2b decision record (why chaining, why not VPC CNI or CNI replacement)02-decision.md, 01-vpccni-encryption-options.log
L2b Cilium install and engine recreate10-cilium-install.log, 11-recreate-engine-under-cilium.log
L2b WireGuard status and peers12-cilium-encryption-status.log, 16-wg-udp-socket-and-cni.log
L2b counter proof (50 MiB transfer)13-wg-counter-proof.log
L2b counters around the bench14-wg-counters-before-bench.log, 15-wg-counters-after-bench.log
L2b benchmark JSON (three runs)bench-L0-before, bench-L2b, bench-L0-after
L2b comparison table and outcome note50-wireguard-tax.md, 60-encryption-proof-and-outcome.md
L2b Cilium-no-encryption attempt (discarded)17-cilium-disable-wg.log, 18-recreate-engine-cilium-noenc.log
Stale CNI conflist after helm uninstall19-cilium-uninstall.log, 22-stale-cni-conflist-breakage.log, 22-remove-stale-cilium-conflist.log
L0-after bracket state23-cilium-reinstall-noenc.log, 24-bracket-state.log
L2b clean uninstall, CNI sanity, L1 restore30-cilium-uninstall-final.log, 31-cni-sanity-pod.log, 32-restore-L1-netpol.log
GPU teardownstageD-99-gpu-teardown; L2b session: 23-tf-plan-gpu0.log, 24-tf-apply-gpu0.log, 25-gpu-zero-verify.log, 40-tf-plan-gpu0.log, 41-gpu-zero-verify.log

Appendix C: Data consistency notes

Discrepancies found while checking the findings summaries against the raw logs. The report uses the raw values.

ItemSummary saidRaw log showsEffect
Stage D L0"re-confirm the distributed baseline" (study plan); L0 listed as "measured"Byte-identical copy of Stage C results, no run logsDisclosed in Baseline provenance (important) and Limitations
L1 at c=32Findings table showed +32.5% throughputCorrect arithmetic, but attributable to the L0 stall, not to L1Reframed as no measurable difference
Stage C worker portsFindings table lists only 10002 on worker as openMatches; the L1 probe also marks non-listening worker ports as BLOCKEDOnly ports open at L0 counted as L1 evidence
L2a proxy behaviourFindings describe L7 503 and opaque-mode resetResult files contain only the 503 run; opaque-mode behaviour is in the outcome note, not in captured proxy logsOpaque-mode claim labelled as operator note
L3 auth proof4 calls listedLog has 6 lines, including one correct-key completion that timed out (curl exit 28) before a successful retryReported both
Stage B /metrics families6612-metrics-names.log counts 66 # HELP linesConsistent
Stage A findings "7 ephemeral RPC ports" on head7nmap shows 7 high ports (36087, 36349, 44217, 44227, 52365, 57914, 62226)Consistent
Checkov versionnot statednot loggedStated as unknown
Linkerd version"edge"exact version not loggedStated as unknown
L2b comparison basisFindings first reported L2b vs L0-before only (-6.2% c=16, -5.3% c=32 throughput)Both bases are correct arithmetic from the JSON; vs mean of both L0 runs gives -4.5% and -6.5%The report uses the mean of both L0 runs; L0-before-only values are in 50-wireguard-tax.md
L2b rangeFindings said "about 4 to 6% throughput, 4 to 8% e2e"vs L0 mean: 3.4% to 6.5% throughput, 3.2% to 8.4% E2E p50The report gives 3% to 8% overall with per-level values
L2b L0-after stateOutcome note says pods on plain VPC CNI pathHead pod confirmed not a Cilium endpoint; worker not checkedStated as head-only evidence
Stage C GPU peaks"head ~13%, worker ~68%"head max 15%, median 12%; worker max 68%, median 37%The report gives min, median, max
Raw per-request data (added 2026-09-29)Tables implied distribution-level analysis was possiblebench/bench.py line 119 writes {"summary","raw"} inside the pod; manifests/stageB-bench-job.yaml line 49 prints only ['summary']; every archived result-c*.json has only p50/p95/p99/meanStated in Benchmark harness and Limitations; no distributions or CIs reported
L1 proof port count (added 2026-09-29)Summary of layers said "6 open ports to blocked"7 ports open at L0 went to BLOCKED: head 6379 and 10002 to 10006 (6), worker 10002 (1)Corrected to 7
L1 scope (added 2026-09-29)"L1 closes the cross-namespace exposure"Every policy is policyTypes: ["Ingress"] (line 22 for default-deny); egress deliberately open (comment at line 11)Stated as ingress-only, no containment claim
L3 scope (added 2026-09-29)"removed unauthenticated inference"01-auth-proof.log covers /v1/models and /v1/completions only; v0.11.0 is in the CVE-2026-48746 rangeNarrowed to two /v1 routes; CVE added
L2a wording (added 2026-09-29)"incompatible"Only the 503 result files exist; no proxy logs, no Linkerd versionReworded to "tested configuration failed; root cause not isolated"
L0-after bracket (added 2026-09-29)"valid no-encryption L0", "plain VPC CNI path"24-bracket-state.log: Cilium agent running, encryption off, no cilium_wg0; head not an endpoint; worker unchecked; final uninstall at 13:22:22ZStated as not a proven Cilium-free baseline
Actor for rename and scale-down (added 2026-09-29)attributed to a different session actorSame study team, parallel control sessionWording corrected here and in 60-encryption-proof-and-outcome.md
VPC CNI enforcing mode (added 2026-09-29)not discussed01-vpccni-encryption-options.log lists NETWORK_POLICY_ENFORCING_MODE as a variable name only, no valueStartup window stated as unmeasured

Appendix D: Response to independent review (2026-09-29)

Corrections after two independent reviews were applied on 2026-09-29. This appendix and Appendix E record them.

Two independent reviews of the draft were received on 2026-09-29. One review was received as a secondary copy whose size and hash did not match what the primary review described, so we treated it as a secondary copy of the review points and checked every point against the raw logs and harness ourselves. External facts were re-verified on the web on 2026-09-29 (references [4], [18], [19]).

Review pointDecisionWhere fixed
API key "removed unauthenticated inference" is too strong; only two /v1 routes were testedAcceptedSummary, Contribution 4, Reachability matrix, Security effect of each layer, Summary of layers, Practical recommendations, Defence in depth ordering, Limitations, Conclusion
vLLM v0.11.0 is in the CVE-2026-48746 / GHSA-94f4-hr76-p5j6 affected rangeAccepted (verified: >=0.3.0 <0.22.0, published 2026-06-02; bypass not tested here)Distributed LLM inference, Software under test, Security effect of each layer, Related Work, References [19]
Ray token authentication and TLS were omittedAccepted (verified: token auth from 2.52.0, off by default; not evaluated). Reference [4] upgraded from [U] to [C]Kubernetes networking and NetworkPolicy, Software under test, Practical recommendations, Defence in depth ordering, Limitations, Related Work, Future work
WireGuard figure is Cilium chaining plus WireGuard, n=1Accepted; already disclosed, wording tightened to "associated with"Summary, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Network-layer versus L7 encryption for streaming inference, Conclusion
Linkerd "incompatible" is not demonstratedAcceptedContribution 5, L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload), Summary of layers, Network-layer versus L7 encryption for streaming inference, Related Work, Ethics and Responsible Disclosure, Conclusion
Per-request raw samples were never persistedAccepted (verified: bench/bench.py line 119 vs manifests/stageB-bench-job.yaml line 49)Benchmark harness, Limitations, Appendix A, Appendix C
L1 is ingress-only; no containment evidenceAccepted (verified: line 22 policyTypes: ["Ingress"], line 11 egress left open)Kubernetes networking and NetworkPolicy, Hardening layers, Security effect of each layer, Summary of layers, Practical recommendations, Defence in depth ordering, Limitations
containerPort is metadata, not a firewallAcceptedSummary, Contribution 2, Kubernetes networking and NetworkPolicy, Four default scanner configurations and the runtime surface, What the four default scanner configurations did not report, and why
VPC CNI standard mode has a default-allow startup windowPartly accepted: the documented behaviour is real, but no evidence either way for this study, because the value was not recorded. Presented as a gap, not a findingKubernetes networking and NetworkPolicy, Environment, Reachability matrix, Security effect of each layer, Limitations
Cilium removal failure should be framed as incomplete procedure with a documented defaultAcceptedSummary, Contribution 7, L2b: Cilium chaining plus WireGuard combined (single bracketed run)
L0 reused for L1 and L3 is cross-sessionAccepted; already disclosed in 4.6no change
Tag-based images; digests and some versions not recordedAcceptedSoftware under test, Limitations, Future work
Wilson lower bound for 128 of 128 successes is about 97%Accepted as a note; not added to tables, since success counts are not a headline claimnone
Proposed designs: 10 blocks, 256 measured requests per run, randomisation seed 20260929Partly accepted: the 10-block minimum is adopted in Future work; the specific request count and seed are the reviewers' proposal, not requirementsFuture work
Cost per million tokens and a Pareto frontierPartly accepted as a later extension; not in this report's scopenot added
Infrastructure-agent misalignment study and telemetry-assisted prefix-cache leakageRejected for this report: outside the research questions. May be separate studiesnot added
"Universal cluster-wide reachability" should not be claimedAccepted; the report claims reach from one neighbour pod in one testbedConclusion

Items from the reviews that were already correct in the report (L0 reuse, counter-based rather than pcap WireGuard evidence, L2b drift) were left as they were.

Appendix E: Response to a second independent review (2026-09-29)

A second independent review was received on 2026-09-29. It rated the cross-namespace Ray exposure, the cleartext Ray control-plane evidence and the NetworkPolicy blocking as the strong results, and the performance numbers and generalisation as the weakest. We checked each point against the report and the stored logs. No experiment was run and no measured number was changed. Only wording, framing and structure changed, plus an offline re-reading of the stored scanner outputs.

Review pointDecisionWhat changedWhere
Headline overstated: one constructed neighbour pod, one clusterAccepted. The old conclusion generalised the result to the whole cluster, which the test did not showTitle, subtitle, abstract and conclusion now state the constrained result: a pod in an unrelated namespace could reach Ray control-plane listeners under these defaults, in one cluster, one topology and one version setTitle, Summary, Conclusion
Scanners: the claim could be read as a general oneAcceptedRestated as four default scanner configurations that did not identify the RayCluster's runtime network exposure. Four default scanner configurations and the runtime surface says these are runtime and network properties and that custom rules were not tested. Offline re-analysis of the stored outputs: no output names the RayCluster, and rules of the kind that fired on the two Deployments would plausibly apply to its pod templates. No rule concerns reachability. The 57-versus-zero comparison is not like for likeSummary, Introduction, Four default scanner configurations and the runtime surface, What the four default scanner configurations did not report, and why, Future work
containerPort argument is dramaticAcceptedReframed as visibility drift for tooling that treats declared ports as an exposure inventory. containerPort is metadata, not a security boundarySummary, Introduction, Four default scanner configurations and the runtime surface
Performance called a "cost" from n=1Accepted. The numbers are unchanged"Cost" and "tax" became "observed difference in this single bracketed run" in prose, headings and table headers. The limits are stated next to the figuresSummary, Introduction, Performance comparison, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion, Limitations, Conclusion
WireGuard confoundAccepted; the confound was already disclosedThe configuration is labelled Cilium chaining plus WireGuard combined throughout, and the text says WireGuard alone was not isolated. The A, B and C design was already in Future workSummary, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion, Future work, Conclusion
API key result given too much weightAccepted; the caveats were already presentMoved from a headline result to a weak one. It shows the normal key path returned 401 and does not show that API-key authentication protects vLLM. Recommendation order changedSummary, scope box, Security effect of each layer, Practical recommendations, Defence in depth ordering, Conclusion
Ray RCE could be misreadAcceptedAdded a claim-by-claim table. Job API can be exposed: yes. Default CPU Ray exposed it: yes. GPU vLLM exposed it: no. Remote code execution against the GPU inference stack: no. Ray control ports reachable cross-namespace: yesScope box, Summary, The distributed control plane: cleartext and reachable
NCCL encryption stated as factAcceptedWorded as inference from the WireGuard datapath carrying inter-node pod traffic. NCCL traffic was not separately identified or capturedSummary, The distributed control plane: cleartext and reachable, L2b: Cilium chaining plus WireGuard combined (single bracketed run), Summary of layers, Discussion
Startup window unansweredAccepted. A search of all stored logs found the variable name only in an add-on option listing, with no valueRecorded as not recorded and as an open question: newly created Ray pods during startup were not testedLimitations, Future work
Linkerd is an incident, not a security findingAcceptedLabelled an engineering incident and moved out of the main conclusions and takeaways. The data stays in L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload)Summary, Introduction, L2a: Linkerd mTLS (engineering incident: the tested configuration failed to carry the workload), Conclusion

Left as it was: the L0 reuse disclosure, the counter-based encryption evidence, the drift analysis and all measured values. The reviewer's three open questions (what the exposed Ray ports allow beyond reachability, whether NetworkPolicy is still needed with Ray native authentication or TLS, and whether the scanner comparison is fair) would need new experiments or new scans. They are recorded in Future work as not tested.

Third round of wording changes (2026-09-29, no experiment, no measured number changed): generalisations beyond the tested KubeRay and vLLM topology were removed, the scanner result now leads with the RayCluster not being inspected (finding counts stay in the Four default scanner configurations and the runtime surface table), the Ray Job API recommendation now reads as a potential arbitrary-code-execution interface resting on documented semantics, a Demonstrated / Not demonstrated / Engineering incidents summary was added after the scope box, and Future work now names rendering the RayCluster into pods and rerunning the same scanners.

About Sorami

Sorami is an Australian cyber security and cloud consultancy. We build and secure cloud environments, Kubernetes and AI agents, and the senior engineers who scope the work deliver it themselves. To discuss this research or a review of your own stack, use the form below. Only your email is required, and we reply within one business day.

An enquiry, not a booking. We use your details only to reply. See our privacy notice, or go to the contact page.

Last reviewed: