A follow-up on running OpenShift AI across a three-node SNUC EE2300 cluster — this time focused on how models actually get served.

Reference: https://github.com/explicitworkload/snuc-openshift-ai – setup instructions are in the git repository.
My thanks to Andy Penman at SNUC for partnering with Red Hat to make this entire setup possible. Having real hardware to push against — rather than a cloud instance that hides every constraint behind an abstraction — is what made the findings below possible. Several of them only surface when you own the silicon.
The cluster
Three SNUC mini PCs, each identical on paper:
| SNUC EE2300 | Per node |
|---|---|
| CPU | AMD Ryzen AI 9 HX 370, 24 vCPU |
| RAM | 96GB RAM |
| GPU | Radeon 890M integrated, gfx1150 |
| Platform | OpenShift 4.21, OpenShift AI 3.5, RHEL CoreOS 9.6 |
No discrete GPUs. One integrated GPU per node, three in total, each sharing system memory with the host. That last detail drives nearly every decision that follows.
The Architecture
Two independent inference paths sit on one serving layer. Vision, which is latency-sensitive and moves binary tensors:
┌───────────────┐
│ RTSP camera │───┐
└───────────────┘ │ ┌─────────────────────┐ ┌──────────────────────────┐
├──▶│ vision-ai │── gRPC :8001 ─▶│ yolo26 detection│
┌───────────────┐ │ │ OpenCV capture + │ KServe v2 │ yolo26-seg + masks│
│ video upload │───┘ │ annotation render │───────────────▶│ runtime: onnxruntime-cpu │
└───────────────┘ └──────────┬──────────┘ └──────────────────────────┘
│
▼
annotated MJPEG · snapshots · JSON
And language, which is streaming text at conversational frequency:
┌───────────────┐ ┌─────────────────────┐ ┌──────────────────────────┐
│ browser │──────▶│ command-center │── HTTP /v1/ ──▶│ llama-4 (GGUF)│
└───────────────┘ │ proxy + discovery │ OpenAI-compat │ runtime: llama.cpp │
│ (ISVCs via k8s API) │ │ Vulkan offload │
└──────────┬──────────┘ └──────────────────────────┘
│
│ ┌──────────────────────────┐
└──────────────────────────▶│ litellm → external │
│ providers │
└──────────────────────────┘
Both terminate in KServe InferenceServices, but they speak different protocols and carry very different request profiles — and that difference drives most of what follows.
Vision: KServe v2 over gRPC
The app captures RTSP frames on a background thread and calls the predictors with tritonclient.grpc — KServe’s v2 inference protocol is Triton-compatible:
inputs = [grpcclient.InferInput("images", blob.shape, "FP32")]
outputs = [grpcclient.InferRequestedOutput("output0")] # boxes
outputs.append(grpcclient.InferRequestedOutput("output1")) # masks, seg only
Detection and segmentation are separate InferenceServices behind the same image, differing only in the ONNX file they load and whether they return output1.
Language: OpenAI-compatible over HTTP
The GGUF runtimes expose /v1/chat/completions directly, so Command Center proxies them without translation, and LiteLLM fronts external providers behind the same shape. One client, one protocol, regardless of whether the model is in-cluster.
Discovery is cluster-native
Command Center has no hardcoded model list. It queries the Kubernetes API for InferenceServices:
custom_api.list_namespaced_custom_object(..., plural="inferenceservices")
and merges that with external endpoints declared in a ConfigMap. Deploy an InferenceService and it appears in the UI; delete it and it disappears. The cluster is the registry, which removes an entire class of config drift.
Bypassing the auth sidecar on the hot path
This one is a genuine tradeoff. KServe fronts each predictor with a kube-rbac-proxy sidecar — that’s why predictor pods run 2/2 — and with enable-auth set, every request carries an authorization check.
For a chat completion a few times a minute, that’s free. For video inference at frame rate, it’s a per-frame tax on a path that never leaves the cluster. So the vision models are reached through plain Services that select the predictor pods directly and target the raw gRPC port:
kind: Service
metadata: {name: yolo26-grpc}
spec:
selector: {app: isvc.yolo26-predictor} # the pod, not the KServe service
ports: [{port: 8001, targetPort: 8001}]
The app is pointed at these:
INFERENCE_ENDPOINTS=yolo26=yolo26-grpc.john.svc.cluster.local:8001,...
Worth being explicit about what this costs: that path has no authorization. It’s defensible for pod-to-pod traffic inside a namespace, and it would not be defensible on anything reachable from outside. The authenticated route still exists and still works — this is a deliberate second door, not a replacement.
Weights live in S3, never in images
KServe’s storage-initializer init container pulls model weights from S3 into /mnt/models before the serving container starts. The image ships the engine; the InferenceService supplies the model.
predictor pod
┌──────────────────────┐ ┌──────────────────────────────────────┐
│ S3 / ODF bucket │ │ ① storage-initializer (init) │
│ │──▶│ pulls weights ──▶ /mnt/models │
│ /models/yolo26m.onnx │ │ ────────────────────────────────── │
│ /models/*-seg.onnx │ │ ② kserve-container │
│ /models/*.gguf │ │ loads from /mnt/models │
└──────────────────────┘ │ ③ kube-rbac-proxy (sidecar) │
secret: └──────────────────────────────────────┘
models-odf-s3 image carries no weights
That’s why one visionai:cpu image serves both vision models, and why swapping a model is a path change rather than a rebuild. It also means image CI and model updates are fully decoupled — a retrained ONNX file needs no pipeline run at all. It also explains the 2/2 you see on predictor pods: the serving container plus the auth sidecar from the previous section.
Namespace topology
Four namespaces, split by lifecycle rather than by application:
| Namespace | Holds | Why separate |
|---|---|---|
john | models, apps, routes | the workload blast radius |
visionai | Tekton pipelines, triggers, webhook route | CI identity and RBAC differ from runtime |
redhat-ods-applications | hardware profiles | platform-owned, shared by all namespaces |
openshift-gitops | ArgoCD Applications | delivery control plane |
Delivery: ArgoCD for state, Tekton for artifacts
The split is clean. Tekton builds images and restarts rollouts; ArgoCD owns declared state and syncs each component’s k8s/ directory. Neither does the other’s job.
One deliberate exception is worth calling out, because it’s the kind of thing that bites later: the hardware profiles live outside any ArgoCD-synced directory. They target a platform namespace, and the vision Application syncs with prune: true — so folding them in would hand ArgoCD authority to delete shared platform objects if a file were ever moved or renamed. They’re applied manually, on purpose, and the manifests say so in a comment.
How a model actually gets served
OpenShift AI serves models through KServe, and it’s worth being precise about the split, because this is where most of the confusion lives.
A ServingRuntime is the engine. It declares a container image, the model formats it understands, and the resources it needs:
kind: ServingRuntime
metadata:
name: onnxruntime-cpu
spec:
supportedModelFormats:
- name: onnx
containers:
- name: kserve-container
image: <registry>/visionai:cpu
args:
- --model_path=/mnt/models
- --providers=CPUExecutionProvider
resources:
requests: {cpu: "1", memory: 2Gi}
An InferenceService is one model served by that engine. It names the runtime, points at weights in S3, and may override resources:
kind: InferenceService
metadata:
name: yolo26
annotations:
opendatahub.io/hardware-profile-name: cpu-only
spec:
predictor:
model:
modelFormat: {name: onnx}
runtime: onnxruntime-cpu
storage: {key: models-odf-s3, path: /models/yolo26m.onnx}
resources:
requests: {cpu: "2", memory: 4Gi}
Two models can share one runtime. Swapping engines is a one-line change. KServe pulls the weights into /mnt/models at startup, so the image carries no model — which is why a single image serves both our detection and segmentation models.
The hardware profile will override you
Here’s the first thing that cost me real time, and the reason I’d put hardware profiles near the top of anyone’s reading list.
OpenShift AI ships exactly one hardware profile, default-profile. On a GPU-equipped cluster it looks like this:
identifiers:
- identifier: amd.com/gpu
resourceType: Accelerator
minCount: 1 # <- not optional
defaultCount: 1
minCount: 1 means every InferenceService referencing that profile reserves a GPU, regardless of whether its runtime can use one. Our two YOLO models were both on a CPU runtime and both holding a GPU they could never touch. Two of three accelerators in the cluster, consumed by models doing pure CPU inference.
The fix is a profile with no accelerator stanza. But there’s a sharp edge: the GPU request is injected into the stored InferenceService spec by an admission webhook. Changing the profile afterwards does not remove it, because kubectl apply merges maps — a key absent from your manifest is left alone, not deleted. You have to strip it explicitly:
oc patch inferenceservice yolo26 -n john --type=json -p \
'[{"op":"remove","path":"/spec/predictor/model/resources/requests/amd.com~1gpu"},
{"op":"remove","path":"/spec/predictor/model/resources/limits/amd.com~1gpu"}]'
Until you do, the manifest says one thing and the cluster does another.
Requests, not utilization
A related trap. One model sat Pending for over two weeks while the cluster looked half-idle — actual CPU utilization was 26–33% on two nodes.
The scheduler does not care about utilization. It schedules on requests, and those were at 81–88% of allocatable CPU. The model asked for 4 cores; the only node with a free GPU had 3886m uncommitted. It missed by 214 millicores — roughly a fifth of a core — and sat there for sixteen hours logging FailedScheduling while the machine it wanted was two-thirds idle.
Dropping the request from 4 to 2 placed it immediately. Over-requesting is not a safety margin on a small cluster; it’s a scheduling failure waiting to happen.
The accelerator story: Vulkan yes, ROCm no
Now the part that matters most if you’re deploying to AMD integrated graphics.
Vulkan works. Our GGUF language models run on llama.cpp built from ghcr.io/ggml-org/llama.cpp:full-vulkan, offloading 99 layers to the iGPU, and they serve happily. Vulkan addresses the shared system memory pool rather than demanding a dedicated VRAM carve-out, which suits an APU.
ROCm does not. ONNX Runtime with ROCMExecutionProvider fails at session creation:
Hip error: 'out of memory'(2) at .../hipblaslt/library/src/amd_detail/hipblaslt.cpp:165
The obvious read is a VRAM problem, and the obvious fix is to raise the BIOS carve-out. We did — from the 512 MB default to 32 GB on two nodes and 48 GB on a third. It failed identically. So I ran the matrix properly, inside the real serving image, on a real Conv model:
| Configuration | Result |
|---|---|
| baseline | FAIL — hipblaslt OOM |
HSA_OVERRIDE_GFX_VERSION=11.0.0 | FAIL — identical |
HSA_OVERRIDE=11.0.0 + ROCBLAS_USE_HIPBLASLT=0 | FAIL — identical |
ROCBLAS_USE_HIPBLASLT=0 | FAIL — identical |
Same error, same source line, every time. Both standard workarounds — masquerading as a supported gfx target, and forcing the rocBLAS fallback — changed nothing. The GPU is visible (/dev/kfd, /dev/dri/renderD128), allocated to the pod, and ROCm 6.4 is installed. It simply cannot initialize hipblaslt on gfx1150.
This is not a VRAM sizing problem, and raising the container memory limit doesn’t help either. Worth knowing before you spend an afternoon on it.
A naming trap worth repeating
Our GPU runtime was called onnxruntime-migraphx and built FROM rocm/migraphx-ci-ubuntu. Reasonable enough. But the image installs the onnxruntime-rocm wheel, which ships only ROCMExecutionProvider and CPUExecutionProvider. There is no MIGraphX provider in it.
So --providers=MIGraphXExecutionProvider,CPUExecutionProvider produced a warning and silently fell through to CPU — while still holding a GPU. The model was named for GPU inference, scheduled as GPU inference, and executing on CPU the entire time. A pip wheel will not give you the MIGraphX EP; that needs ONNX Runtime built from source with --use_migraphx.
Always assert the provider you actually got:
session = ort.InferenceSession(model, providers=[...])
assert "ROCMExecutionProvider" in session.get_providers()
VRAM is not free
One more APU-specific wrinkle. Dedicated VRAM is carved out of system RAM in BIOS, so raising it reduces what pods can schedule. Our 48 GB node dropped to 46 GB of allocatable RAM against 62 GB on its siblings — enough that a 4 GiB pod request stopped fitting. Since ROCm doesn’t work here anyway, that carve-out is pure cost.
Where this leaves us
Both vision models run on CPU: segmentation at roughly 0.7–1.1 s per frame, detection at 3.5–4.8 s. Not real-time, and honest about it. All three GPUs are now unreserved rather than held hostage by models that couldn’t use them.
The realistic path to GPU-accelerated vision on this hardware is not ROCm. It’s an engine with a Vulkan compute backend — ncnn is the obvious candidate, and our own LLM stack already proves Vulkan is sound on these chips. ONNX Runtime has no Vulkan execution provider, so that means changing engines, not flipping a flag.
Takeaways
- ServingRuntime is the engine, InferenceService is the model. Keep weights out of the image and swapping runtimes stays a one-line change.
- Audit your hardware profile.
default-profileforces a GPU reservation withminCount: 1. A CPU model paired with it silently consumes an accelerator. - An injected resource request survives
kubectl apply. Map keys merge. Strip them with a JSON patch or they persist invisibly. - The scheduler reads requests, not utilization. A half-idle cluster can be completely unschedulable.
- Verify the execution provider at runtime. A name in a config proves nothing — ORT will warn and fall back to CPU while you believe you’re on GPU.
- On AMD APUs, reach for Vulkan. It works today where ROCm does not.
Thanks again to Andy Penman and SNUC, and to the Red Hat Singapore team. Constraints make for better engineering, and this cluster supplied plenty.