task e2e-llm-inference-service has failed: "step-fail-if-needed" exited with code 1: Error [get-kubeconfig] Found kubeconfig secret: cluster-kh8hc-admin-kubeconfig [get-kubeconfig] Wrote kubeconfig to /credentials/cluster-kh8hc-kubeconfig [get-kubeconfig] Found admin password secret: cluster-kh8hc-admin-password [get-kubeconfig] Retrieved username [get-kubeconfig] Wrote password to /credentials/cluster-kh8hc-password [get-kubeconfig] API Server URL: https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443 [get-kubeconfig] Console URL: https://console-openshift-console.apps.c60b4ca1-f3c1-464f-bf51-d0e4245a34af.prod.konfluxeaas.com [clone-repo] work/RHOAIENG-79440 [clone-repo] https://github.com/jlost/kserve [clone-repo] Cloning into '/workspace/source'... [clone-repo] Updating files: 88% (2863/3251) Updating files: 89% (2894/3251) Updating files: 90% (2926/3251) Updating files: 91% (2959/3251) Updating files: 92% (2991/3251) Updating files: 93% (3024/3251) Updating files: 94% (3056/3251) Updating files: 95% (3089/3251) Updating files: 96% (3121/3251) Updating files: 97% (3154/3251) Updating files: 98% (3186/3251) Updating files: 99% (3219/3251) Updating files: 100% (3251/3251) Updating files: 100% (3251/3251), done. [e2e-llm-inference-service] + bash [e2e-llm-inference-service] + STATUS_FILE=/test-status/deploy-and-e2e-status [e2e-llm-inference-service] + echo failed [e2e-llm-inference-service] + COMPONENT_NAME=kserve-agent-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-agent-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:440184469bf83caf5543427fdbbf80873e43c9c7af817ffabe7c58eda76f78c2 [e2e-llm-inference-service] + KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:440184469bf83caf5543427fdbbf80873e43c9c7af817ffabe7c58eda76f78c2 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-controller-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-controller-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:a627ec30debc454af106c2c34af695e96eccb25d564ea70da2ca083e0bc41d91 [e2e-llm-inference-service] + KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:a627ec30debc454af106c2c34af695e96eccb25d564ea70da2ca083e0bc41d91 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-router-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-router-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:ecd8623c42b521978afd08b50dfd4d13bfe521fd8131429e4f24493619574031 [e2e-llm-inference-service] + KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:ecd8623c42b521978afd08b50dfd4d13bfe521fd8131429e4f24493619574031 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-storage-initializer-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-storage-initializer-ci '.[$component_name].image' [e2e-llm-inference-service] + export STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] + STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] + COMPONENT_NAME=odh-kserve-llmisvc-controller-ci [e2e-llm-inference-service] ++ jq -r --arg component_name odh-kserve-llmisvc-controller-ci '.[$component_name].image' [e2e-llm-inference-service] + export LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:c5b450eea4db6d5fde1b2fb72cb1a1fa1929d109024104cb2f0ea7314a0fdd04 [e2e-llm-inference-service] + LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:c5b450eea4db6d5fde1b2fb72cb1a1fa1929d109024104cb2f0ea7314a0fdd04 [e2e-llm-inference-service] + ./test/scripts/openshift-ci/run-e2e-tests.sh 'llminferenceservice and cluster_cpu and not autoscaling and not tracing' 2 llm-d [e2e-llm-inference-service] INFERENCE_POOL_GROUP=inference.networking.k8s.io (detected from OCP 4.21.26) [e2e-llm-inference-service] Installing on cluster [e2e-llm-inference-service] Using namespace: kserve for KServe components [e2e-llm-inference-service] SKLEARN_IMAGE=quay.io/opendatahub/sklearn-serving-runtime:odh-pr-1821 [e2e-llm-inference-service] OPT_125M_MODEL_URI=s3://example-models/facebook/opt-125m [e2e-llm-inference-service] ERROR_404_ISVC_IMAGE=quay.io/opendatahub/error-404-isvc:odh-pr-1821 [e2e-llm-inference-service] SUCCESS_200_ISVC_IMAGE=quay.io/opendatahub/success-200-isvc:odh-pr-1821 [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed Kustomize v5.8.1 to /workspace/source/bin/kustomize [e2e-llm-inference-service] v5.8.1 [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] [INFO] Installing yq v4.52.1 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed yq v4.52.1 to /workspace/source/bin/yq [e2e-llm-inference-service] yq (https://github.com/mikefarah/yq/) version v4.52.1 [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Installing KServe Python SDK ... [e2e-llm-inference-service] [INFO] Installing uv 0.7.8 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed uv 0.7.8 to /workspace/source/bin/uv [e2e-llm-inference-service] warning: Failed to read project metadata (No `pyproject.toml` found in current directory or any parent directory). Running `uv self version` for compatibility. This fallback will be removed in the future; pass `--preview` to force an error. [e2e-llm-inference-service] uv 0.7.8 [e2e-llm-inference-service] Creating virtual environment... [e2e-llm-inference-service] warning: virtualenv's `--clear` has no effect (uv always clears the virtual environment) [e2e-llm-inference-service] Using CPython 3.9.25 interpreter at: /usr/bin/python3 [e2e-llm-inference-service] Creating virtual environment at: .venv [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] Using CPython 3.11.13 interpreter at: /usr/bin/python3.11 [e2e-llm-inference-service] Creating virtual environment at: .venv [e2e-llm-inference-service] Resolved 274 packages in 2ms [e2e-llm-inference-service] Building kserve @ file:///workspace/source/python/kserve [e2e-llm-inference-service] Downloading pyarrow (40.1MiB) [e2e-llm-inference-service] Downloading uvloop (3.8MiB) [e2e-llm-inference-service] Downloading aiohttp (1.7MiB) [e2e-llm-inference-service] Downloading kubernetes (1.9MiB) [e2e-llm-inference-service] Downloading portforward (3.9MiB) [e2e-llm-inference-service] Downloading pydantic-core (2.0MiB) [e2e-llm-inference-service] Downloading setuptools (1.2MiB) [e2e-llm-inference-service] Downloading cryptography (4.5MiB) [e2e-llm-inference-service] Downloading pandas (12.5MiB) [e2e-llm-inference-service] Downloading grpcio-tools (2.5MiB) [e2e-llm-inference-service] Downloading numpy (15.7MiB) [e2e-llm-inference-service] Downloading black (1.6MiB) [e2e-llm-inference-service] Downloading grpcio (6.4MiB) [e2e-llm-inference-service] Downloading botocore (12.9MiB) [e2e-llm-inference-service] Downloading mypy (17.2MiB) [e2e-llm-inference-service] Building timeout-sampler==1.0.3 [e2e-llm-inference-service] Building python-simple-logger==2.0.19 [e2e-llm-inference-service] Downloading aiohttp [e2e-llm-inference-service] Downloading pydantic-core [e2e-llm-inference-service] Downloading black [e2e-llm-inference-service] Downloading grpcio-tools [e2e-llm-inference-service] Downloading setuptools [e2e-llm-inference-service] Downloading portforward [e2e-llm-inference-service] Built python-simple-logger==2.0.19 [e2e-llm-inference-service] Downloading uvloop [e2e-llm-inference-service] Downloading cryptography [e2e-llm-inference-service] Downloading grpcio [e2e-llm-inference-service] Downloading kubernetes [e2e-llm-inference-service] Built timeout-sampler==1.0.3 [e2e-llm-inference-service] Downloading numpy [e2e-llm-inference-service] Built kserve @ file:///workspace/source/python/kserve [e2e-llm-inference-service] Downloading pandas [e2e-llm-inference-service] Downloading botocore [e2e-llm-inference-service] Downloading mypy [e2e-llm-inference-service] Downloading pyarrow [e2e-llm-inference-service] Prepared 101 packages in 1.79s [e2e-llm-inference-service] warning: Failed to hardlink files; falling back to full copy. This may lead to degraded performance. [e2e-llm-inference-service] If the cache and target directories are on different filesystems, hardlinking may not be supported. [e2e-llm-inference-service] If this is intentional, set `export UV_LINK_MODE=copy` or use `--link-mode=copy` to suppress this warning. [e2e-llm-inference-service] Installed 101 packages in 286ms [e2e-llm-inference-service] + aiohappyeyeballs==2.6.1 [e2e-llm-inference-service] + aiohttp==3.14.1 [e2e-llm-inference-service] + aiosignal==1.4.0 [e2e-llm-inference-service] + annotated-doc==0.0.4 [e2e-llm-inference-service] + annotated-types==0.7.0 [e2e-llm-inference-service] + anyio==4.9.0 [e2e-llm-inference-service] + attrs==25.3.0 [e2e-llm-inference-service] + avro==1.12.0 [e2e-llm-inference-service] + black==24.3.0 [e2e-llm-inference-service] + boto3==1.37.35 [e2e-llm-inference-service] + botocore==1.37.35 [e2e-llm-inference-service] + cachetools==5.5.2 [e2e-llm-inference-service] + certifi==2025.1.31 [e2e-llm-inference-service] + cffi==2.0.0 [e2e-llm-inference-service] + charset-normalizer==3.4.1 [e2e-llm-inference-service] + click==8.4.2 [e2e-llm-inference-service] + cloudevents==1.11.0 [e2e-llm-inference-service] + colorama==0.4.6 [e2e-llm-inference-service] + colorlog==6.10.1 [e2e-llm-inference-service] + coverage==7.8.0 [e2e-llm-inference-service] + cryptography==49.0.0 [e2e-llm-inference-service] + deprecation==2.1.0 [e2e-llm-inference-service] + durationpy==0.9 [e2e-llm-inference-service] + execnet==2.1.1 [e2e-llm-inference-service] + fastapi==0.136.3 [e2e-llm-inference-service] + frozenlist==1.5.0 [e2e-llm-inference-service] + google-auth==2.39.0 [e2e-llm-inference-service] + grpc-interceptor==0.15.4 [e2e-llm-inference-service] + grpcio==1.78.1 [e2e-llm-inference-service] + grpcio-testing==1.78.1 [e2e-llm-inference-service] + grpcio-tools==1.78.1 [e2e-llm-inference-service] + h11==0.16.0 [e2e-llm-inference-service] + httpcore==1.0.9 [e2e-llm-inference-service] + httptools==0.6.4 [e2e-llm-inference-service] + httpx==0.27.2 [e2e-llm-inference-service] + httpx-retries==0.4.5 [e2e-llm-inference-service] + idna==3.10 [e2e-llm-inference-service] + iniconfig==2.1.0 [e2e-llm-inference-service] + jinja2==3.1.6 [e2e-llm-inference-service] + jmespath==1.0.1 [e2e-llm-inference-service] + kserve==0.20.0rc0 (from file:///workspace/source/python/kserve) [e2e-llm-inference-service] + kubernetes==32.0.1 [e2e-llm-inference-service] + markupsafe==3.0.2 [e2e-llm-inference-service] + multidict==6.4.3 [e2e-llm-inference-service] + mypy==0.991 [e2e-llm-inference-service] + mypy-extensions==1.0.0 [e2e-llm-inference-service] + numpy==2.2.4 [e2e-llm-inference-service] + oauthlib==3.2.2 [e2e-llm-inference-service] + orjson==3.11.9 [e2e-llm-inference-service] + packaging==24.2 [e2e-llm-inference-service] + pandas==2.2.3 [e2e-llm-inference-service] + pathspec==0.12.1 [e2e-llm-inference-service] + platformdirs==4.3.7 [e2e-llm-inference-service] + pluggy==1.5.0 [e2e-llm-inference-service] + portforward==0.7.1 [e2e-llm-inference-service] + prometheus-client==0.21.1 [e2e-llm-inference-service] + propcache==0.3.1 [e2e-llm-inference-service] + protobuf==6.33.5 [e2e-llm-inference-service] + psutil==5.9.8 [e2e-llm-inference-service] + pyarrow==19.0.1 [e2e-llm-inference-service] + pyasn1==0.6.3 [e2e-llm-inference-service] + pyasn1-modules==0.4.2 [e2e-llm-inference-service] + pycparser==2.22 [e2e-llm-inference-service] + pydantic==2.12.4 [e2e-llm-inference-service] + pydantic-core==2.41.5 [e2e-llm-inference-service] + pyjwt==2.13.0 [e2e-llm-inference-service] + pytest==7.4.4 [e2e-llm-inference-service] + pytest-asyncio==0.23.8 [e2e-llm-inference-service] + pytest-cov==5.0.0 [e2e-llm-inference-service] + pytest-httpx==0.30.0 [e2e-llm-inference-service] + pytest-json-report==1.5.0 [e2e-llm-inference-service] + pytest-metadata==3.1.1 [e2e-llm-inference-service] + pytest-xdist==3.6.1 [e2e-llm-inference-service] + python-dateutil==2.9.0.post0 [e2e-llm-inference-service] + python-dotenv==1.1.0 [e2e-llm-inference-service] + python-multipart==0.0.32 [e2e-llm-inference-service] + python-simple-logger==2.0.19 [e2e-llm-inference-service] + pytz==2025.2 [e2e-llm-inference-service] + pyyaml==6.0.2 [e2e-llm-inference-service] + requests==2.32.3 [e2e-llm-inference-service] + requests-oauthlib==2.0.0 [e2e-llm-inference-service] + rsa==4.9.1 [e2e-llm-inference-service] + s3transfer==0.11.4 [e2e-llm-inference-service] + setuptools==78.1.0 [e2e-llm-inference-service] + six==1.17.0 [e2e-llm-inference-service] + sniffio==1.3.1 [e2e-llm-inference-service] + starlette==1.2.1 [e2e-llm-inference-service] + tabulate==0.9.0 [e2e-llm-inference-service] + timeout-sampler==1.0.3 [e2e-llm-inference-service] + timing-asgi==0.3.1 [e2e-llm-inference-service] + tomlkit==0.13.2 [e2e-llm-inference-service] + typing-extensions==4.15.0 [e2e-llm-inference-service] + typing-inspection==0.4.2 [e2e-llm-inference-service] + tzdata==2025.2 [e2e-llm-inference-service] + urllib3==2.7.0 [e2e-llm-inference-service] + uvicorn==0.34.1 [e2e-llm-inference-service] + uvloop==0.21.0 [e2e-llm-inference-service] + watchfiles==1.0.5 [e2e-llm-inference-service] + websocket-client==1.8.0 [e2e-llm-inference-service] + websockets==15.0.1 [e2e-llm-inference-service] + yarl==1.20.0 [e2e-llm-inference-service] Audited 1 package in 47ms [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [INFO] Kustomize v5.8.1 is already installed in /workspace/source/bin (>= v5.8.1) [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Now using project "kserve" on server "https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443". [e2e-llm-inference-service] [e2e-llm-inference-service] You can add applications to this project with the 'new-app' command. For example, try: [e2e-llm-inference-service] [e2e-llm-inference-service] oc new-app rails-postgresql-example [e2e-llm-inference-service] [e2e-llm-inference-service] to build a new example application in Ruby. Or use kubectl to deploy a simple Kubernetes application: [e2e-llm-inference-service] [e2e-llm-inference-service] kubectl create deployment hello-node --image=registry.k8s.io/e2e-test-images/agnhost:2.43 -- /agnhost serve-hostname [e2e-llm-inference-service] [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [INFO] Kustomize v5.8.1 is already installed in /workspace/source/bin (>= v5.8.1) [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Creating namespace openshift-keda... [e2e-llm-inference-service] namespace/openshift-keda created [e2e-llm-inference-service] Namespace openshift-keda created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Creating OperatorGroup openshift-keda... [e2e-llm-inference-service] operatorgroup.operators.coreos.com/openshift-keda created [e2e-llm-inference-service] OperatorGroup openshift-keda created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Creating Subscription for openshift-custom-metrics-autoscaler-operator... [e2e-llm-inference-service] subscription.operators.coreos.com/openshift-custom-metrics-autoscaler-operator created [e2e-llm-inference-service] Subscription openshift-custom-metrics-autoscaler-operator created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Waiting for openshift-custom-metrics-autoscaler-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (0/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (5/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (10/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (15/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (20/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (25/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 found, but not yet Succeeded (Phase: Installing). Waiting... (30/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 found, but not yet Succeeded (Phase: Installing). Waiting... (35/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 found, but not yet Succeeded (Phase: Installing). Waiting... (40/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 found, but not yet Succeeded (Phase: Installing). Waiting... (45/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 found, but not yet Succeeded (Phase: Installing). Waiting... (50/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-2 is ready (Phase: Succeeded). [e2e-llm-inference-service] --- [e2e-llm-inference-service] Applying KedaController custom resource... [e2e-llm-inference-service] Warning: resource kedacontrollers/keda is missing the kubectl.kubernetes.io/last-applied-configuration annotation which is required by oc apply. oc apply should only be used on resources created declaratively by either oc create --save-config or oc apply. The missing annotation will be patched automatically. [e2e-llm-inference-service] kedacontroller.keda.sh/keda configured [e2e-llm-inference-service] KedaController custom resource applied. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Allowing time for KEDA components to be provisioned by the operator ... [e2e-llm-inference-service] Waiting for KEDA Operator pod (selector: "app=keda-operator") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-operator" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-operator" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-operator" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-operator-77f648d694-q6xf7 1/1 Running 0 41s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-operator" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-operator-77f648d694-q6xf7 condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-operator" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Operator pod is ready. [e2e-llm-inference-service] Waiting for KEDA Metrics API Server pod (selector: "app=keda-metrics-apiserver") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-metrics-apiserver" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-metrics-apiserver" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-metrics-apiserver" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-metrics-apiserver-55557df8c-dvhg2 1/1 Running 0 46s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-metrics-apiserver" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-metrics-apiserver-55557df8c-dvhg2 condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-metrics-apiserver" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Metrics API Server pod is ready. [e2e-llm-inference-service] Waiting for KEDA Webhook pod (selector: "app=keda-admission-webhooks") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-admission-webhooks" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-admission-webhooks" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-admission-webhooks" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-admission-56f9798bb4-g8d6q 1/1 Running 0 51s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-admission-webhooks" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-admission-56f9798bb4-g8d6q condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-admission-webhooks" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Webhook pod is ready. [e2e-llm-inference-service] --- [e2e-llm-inference-service] ✅ KEDA deployment script finished successfully. [e2e-llm-inference-service] KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:a627ec30debc454af106c2c34af695e96eccb25d564ea70da2ca083e0bc41d91 [e2e-llm-inference-service] LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:c5b450eea4db6d5fde1b2fb72cb1a1fa1929d109024104cb2f0ea7314a0fdd04 [e2e-llm-inference-service] KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:440184469bf83caf5543427fdbbf80873e43c9c7af817ffabe7c58eda76f78c2 [e2e-llm-inference-service] KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:ecd8623c42b521978afd08b50dfd4d13bfe521fd8131429e4f24493619574031 [e2e-llm-inference-service] STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] Installing KServe via kustomize... [e2e-llm-inference-service] === Final params.env [e2e-llm-inference-service] kserve-controller=quay.io/opendatahub/kserve-controller@sha256:a627ec30debc454af106c2c34af695e96eccb25d564ea70da2ca083e0bc41d91 [e2e-llm-inference-service] llmisvc-controller=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:c5b450eea4db6d5fde1b2fb72cb1a1fa1929d109024104cb2f0ea7314a0fdd04 [e2e-llm-inference-service] kserve-agent=quay.io/opendatahub/kserve-agent@sha256:440184469bf83caf5543427fdbbf80873e43c9c7af817ffabe7c58eda76f78c2 [e2e-llm-inference-service] kserve-router=quay.io/opendatahub/kserve-router@sha256:ecd8623c42b521978afd08b50dfd4d13bfe521fd8131429e4f24493619574031 [e2e-llm-inference-service] kserve-storage-initializer=quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] kserve-llm-d=registry.redhat.io/rhaiis/vllm-cuda-rhel9@sha256:fc68d623d1bfc36c8cb2fe4a71f19c8578cfb420ce8ce07b20a02c1ee0be0cf3 [e2e-llm-inference-service] kserve-llm-d-inference-scheduler=quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] kserve-llm-d-routing-sidecar=quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0 [e2e-llm-inference-service] kserve-llm-d-uds-tokenizer=quay.io/opendatahub/llm-d-kv-cache:v0.8.0 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda=registry.redhat.io/rhaiis/vllm-cuda-rhel9@sha256:fc68d623d1bfc36c8cb2fe4a71f19c8578cfb420ce8ce07b20a02c1ee0be0cf3 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda-fast-1=registry.redhat.io/rhaiis/vllm-cuda-rhel9@sha256:fc68d623d1bfc36c8cb2fe4a71f19c8578cfb420ce8ce07b20a02c1ee0be0cf3 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda-fast-2=registry.redhat.io/rhaiis/vllm-cuda-rhel9@sha256:fc68d623d1bfc36c8cb2fe4a71f19c8578cfb420ce8ce07b20a02c1ee0be0cf3 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda-fast-1-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-nvidia-cuda-fast-2-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-amd-rocm=registry.redhat.io/rhaiis/vllm-rocm-rhel9@sha256:d9a48add238cc095fa43eeee17c8c4d104de60c4dc623e0bc7f8c4b53b2b2e97 [e2e-llm-inference-service] kserve-llm-d-amd-rocm-fast-1=registry.redhat.io/rhaiis/vllm-rocm-rhel9@sha256:d9a48add238cc095fa43eeee17c8c4d104de60c4dc623e0bc7f8c4b53b2b2e97 [e2e-llm-inference-service] kserve-llm-d-amd-rocm-fast-2=registry.redhat.io/rhaiis/vllm-rocm-rhel9@sha256:d9a48add238cc095fa43eeee17c8c4d104de60c4dc623e0bc7f8c4b53b2b2e97 [e2e-llm-inference-service] kserve-llm-d-amd-rocm-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-amd-rocm-fast-1-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-amd-rocm-fast-2-upstream-version=0.11.0+rhai5 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi=registry.redhat.io/rhaii-early-access/vllm-gaudi-rhel9:3.4.0-ea.2 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi-fast-1=registry.redhat.io/rhaii-early-access/vllm-gaudi-rhel9:3.4.0-ea.2 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi-fast-2=registry.redhat.io/rhaii-early-access/vllm-gaudi-rhel9:3.4.0-ea.2 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi-upstream-version=0.16.0 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi-fast-1-upstream-version=0.16.0 [e2e-llm-inference-service] kserve-llm-d-intel-gaudi-fast-2-upstream-version=0.16.0 [e2e-llm-inference-service] kserve-llm-d-ibm-spyre=registry.redhat.io/rhaiis/vllm-spyre-rhel9@sha256:80ae3e435a5be2c1f117f36599103ab05357917dd6e37f0df6613cb3ac2c13ea [e2e-llm-inference-service] kserve-llm-d-ibm-spyre-fast-1=registry.redhat.io/rhaiis/vllm-spyre-rhel9@sha256:80ae3e435a5be2c1f117f36599103ab05357917dd6e37f0df6613cb3ac2c13ea [e2e-llm-inference-service] kserve-llm-d-ibm-spyre-fast-2=registry.redhat.io/rhaiis/vllm-spyre-rhel9@sha256:80ae3e435a5be2c1f117f36599103ab05357917dd6e37f0df6613cb3ac2c13ea [e2e-llm-inference-service] kserve-llm-d-ibm-spyre-upstream-version=0.10.2 [e2e-llm-inference-service] kserve-llm-d-ibm-spyre-fast-1-upstream-version=0.10.2 [e2e-llm-inference-service] kserve-llm-d-ibm-spyre-fast-2-upstream-version=0.10.2 [e2e-llm-inference-service] kserve-llm-d-latency-predictor-prediction=quay.io/opendatahub/odh-latency-predictor-prediction:odh-stable [e2e-llm-inference-service] kserve-llm-d-latency-predictor-training=quay.io/opendatahub/odh-latency-predictor-training:odh-stable [e2e-llm-inference-service] # TODO update when our changes are introduced in the official image [e2e-llm-inference-service] kube-rbac-proxy=quay.io/opendatahub/odh-kube-auth-proxy@sha256:dcb09fbabd8811f0956ef612a0c9ddd5236804b9bd6548a0647d2b531c9d01b3 [e2e-llm-inference-service] kserve-localmodel-controller=quay.io/opendatahub/odh-kserve-localmodel-controller:odh-master [e2e-llm-inference-service] kserve-localmodelnode-agent=quay.io/opendatahub/odh-kserve-localmodelnode-agent:odh-master [e2e-llm-inference-service] ovms-versioning-ubi-micro=registry.redhat.io/ubi9/ubi-micro@sha256:38e934147827349f2b8b11ac9c38d7be23cb8e29f128b1c46e5c7ae54a2d23cd [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/dscinitializations.dscinitialization.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencemodelrewrites.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceobjectives.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] Waiting for CRDs to be established... [e2e-llm-inference-service] Waiting for CRD inferenceservices.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD inferenceservices.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD llminferenceserviceconfigs.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD llminferenceserviceconfigs.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD clusterstoragecontainers.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD clusterstoragecontainers.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD datascienceclusters.datasciencecluster.opendatahub.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD datascienceclusters.datasciencecluster.opendatahub.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io condition met [e2e-llm-inference-service] Applying all resources... [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/dscinitializations.dscinitialization.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencemodelrewrites.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceobjectives.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] serviceaccount/kserve-controller-manager serverside-applied [e2e-llm-inference-service] serviceaccount/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-leader-election-role serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/llmisvc-leader-election-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-admin serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-edit serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-inferenceservice-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-metrics-reader-cluster-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-admin serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-edit serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-view serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-proxy-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-tls-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-view serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/llmisvc-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-inferenceservice-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-llmisvc-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-manager-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-proxy-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-tls-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/llmisvc-manager-rolebinding serverside-applied [e2e-llm-inference-service] configmap/inferenceservice-config serverside-applied [e2e-llm-inference-service] configmap/kserve-parameters serverside-applied [e2e-llm-inference-service] secret/kserve-webhook-server-secret serverside-applied [e2e-llm-inference-service] secret/mlpipeline-s3-artifact serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-metrics-service serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-service serverside-applied [e2e-llm-inference-service] service/kserve-webhook-server-service serverside-applied [e2e-llm-inference-service] service/llmisvc-controller-manager-service serverside-applied [e2e-llm-inference-service] service/llmisvc-webhook-server-service serverside-applied [e2e-llm-inference-service] service/s3-service serverside-applied [e2e-llm-inference-service] deployment.apps/kserve-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/seaweedfs serverside-applied [e2e-llm-inference-service] networkpolicy.networking.k8s.io/kserve-controller-manager serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] clusterstoragecontainer.serving.kserve.io/default serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-router-route serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler-latency-predictor serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tokenizer serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tracing serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-worker-data-parallel serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferencegraph.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceserviceconfig.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/servingruntime.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/trainedmodel.serving.kserve.io serverside-applied [e2e-llm-inference-service] Waiting for llmisvc-controller-manager to be ready... [e2e-llm-inference-service] Waiting for pod -l "control-plane=llmisvc-controller-manager" in namespace "kserve" to be created... [e2e-llm-inference-service] Pod -l "control-plane=llmisvc-controller-manager" in namespace "kserve" found. [e2e-llm-inference-service] Current pods for -l "control-plane=llmisvc-controller-manager" in namespace "kserve": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] llmisvc-controller-manager-5f9554fc9c-4v2sd 0/1 Running 0 6s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=llmisvc-controller-manager" in namespace "kserve" to become ready... [e2e-llm-inference-service] pod/llmisvc-controller-manager-5f9554fc9c-4v2sd condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=llmisvc-controller-manager" in namespace "kserve" are ready. [e2e-llm-inference-service] Re-Applying all resources... [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/dscinitializations.dscinitialization.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencemodelrewrites.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceobjectives.llm-d.ai serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] serviceaccount/kserve-controller-manager serverside-applied [e2e-llm-inference-service] serviceaccount/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-leader-election-role serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/llmisvc-leader-election-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-admin serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-edit serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-inferenceservice-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-metrics-reader-cluster-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-admin serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-edit serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-models-view serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-proxy-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-tls-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-view serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/llmisvc-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-inferenceservice-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-llmisvc-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-manager-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-proxy-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-tls-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/llmisvc-manager-rolebinding serverside-applied [e2e-llm-inference-service] configmap/inferenceservice-config serverside-applied [e2e-llm-inference-service] configmap/kserve-parameters serverside-applied [e2e-llm-inference-service] secret/kserve-webhook-server-secret serverside-applied [e2e-llm-inference-service] secret/mlpipeline-s3-artifact serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-metrics-service serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-service serverside-applied [e2e-llm-inference-service] service/kserve-webhook-server-service serverside-applied [e2e-llm-inference-service] service/llmisvc-controller-manager-service serverside-applied [e2e-llm-inference-service] service/llmisvc-webhook-server-service serverside-applied [e2e-llm-inference-service] service/s3-service serverside-applied [e2e-llm-inference-service] deployment.apps/kserve-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/seaweedfs serverside-applied [e2e-llm-inference-service] networkpolicy.networking.k8s.io/kserve-controller-manager serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] clusterstoragecontainer.serving.kserve.io/default serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-pd-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-multi-node-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-router-route serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler-latency-predictor serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-pd-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-single-node-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tokenizer serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tracing serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-worker-data-parallel serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferencegraph.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceserviceconfig.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/servingruntime.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/trainedmodel.serving.kserve.io serverside-applied [e2e-llm-inference-service] Applying DSC/DSCI resources... [e2e-llm-inference-service] dscinitialization.dscinitialization.opendatahub.io/test-dsci created [e2e-llm-inference-service] datasciencecluster.datasciencecluster.opendatahub.io/test-dsc created [e2e-llm-inference-service] KServe manual installation complete [e2e-llm-inference-service] 🔧 Configuration: [e2e-llm-inference-service] KServe deployment: ❌ disabled [e2e-llm-inference-service] Kuadrant deployment: ✅ enabled [e2e-llm-inference-service] [e2e-llm-inference-service] Checking OpenShift server version...(4.21.26) [e2e-llm-inference-service] 🎯 Server version (4.21.26) is 4.19.9 or higher - continue with the script [e2e-llm-inference-service] ⏳ Installing cert-manager [e2e-llm-inference-service] namespace/cert-manager-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/openshift-cert-manager-operator created [e2e-llm-inference-service] subscription.operators.coreos.com/openshift-cert-manager-operator created [e2e-llm-inference-service] Waiting for openshift-cert-manager-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (0/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (5/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (10/300) [e2e-llm-inference-service] CSV cert-manager-operator.v1.20.0 found, but not yet Succeeded (Phase: Installing). Waiting... (15/300) [e2e-llm-inference-service] CSV cert-manager-operator.v1.20.0 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD certificates.cert-manager.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD certificates.cert-manager.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/certificates.cert-manager.io condition met [e2e-llm-inference-service] ✅ Cert-manager installed [e2e-llm-inference-service] ⏳ Installing openshift-lws-operator [e2e-llm-inference-service] namespace/openshift-lws-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/leader-worker-set created [e2e-llm-inference-service] subscription.operators.coreos.com/leader-worker-set created [e2e-llm-inference-service] Waiting for leader-worker-set CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (0/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (5/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (10/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (15/300) [e2e-llm-inference-service] CSV leader-worker-set.v1.0.0 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD leaderworkersetoperators.operator.openshift.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD leaderworkersetoperators.operator.openshift.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/leaderworkersetoperators.operator.openshift.io condition met [e2e-llm-inference-service] leaderworkersetoperator.operator.openshift.io/cluster created [e2e-llm-inference-service] ⏳ waiting for openshift-lws-operator to be ready.… [e2e-llm-inference-service] Waiting for pod -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" to be created... [e2e-llm-inference-service] Pod -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" found. [e2e-llm-inference-service] Current pods for -l "name=openshift-lws-operator" in namespace "openshift-lws-operator": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] openshift-lws-operator-fd8ccff4c-fxmrs 1/1 Running 0 12s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" to become ready... [e2e-llm-inference-service] pod/openshift-lws-operator-fd8ccff4c-fxmrs condition met [e2e-llm-inference-service] Pod(s) -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" are ready. [e2e-llm-inference-service] ✅ openshift-lws-operator installed [e2e-llm-inference-service] gatewayclass.gateway.networking.k8s.io/openshift-default created [e2e-llm-inference-service] Waiting for pod -l "app=istiod" in namespace "openshift-ingress" to be created... [e2e-llm-inference-service] Pod -l "app=istiod" in namespace "openshift-ingress" found. [e2e-llm-inference-service] Current pods for -l "app=istiod" in namespace "openshift-ingress": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] istiod-openshift-gateway-5677ccb49b-jpnp5 1/1 Running 0 6s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "app=istiod" in namespace "openshift-ingress" to become ready... [e2e-llm-inference-service] pod/istiod-openshift-gateway-5677ccb49b-jpnp5 condition met [e2e-llm-inference-service] Pod(s) -l "app=istiod" in namespace "openshift-ingress" are ready. [e2e-llm-inference-service] ⏳ Creating a Gateway [e2e-llm-inference-service] Error from server (AlreadyExists): namespaces "openshift-ingress" already exists [e2e-llm-inference-service] ⏳ Creating gateway memory ConfigMap for parametersRef (2Gi) [e2e-llm-inference-service] configmap/gateway-proxy-config created [e2e-llm-inference-service] gateway.gateway.networking.k8s.io/openshift-ai-inference created [e2e-llm-inference-service] Waiting for pod -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" to be created... [e2e-llm-inference-service] Pod -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" found. [e2e-llm-inference-service] Current pods for -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] openshift-ai-inference-openshift-default-8699b59bbc-4gqnw 1/1 Running 0 6s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" to become ready... [e2e-llm-inference-service] pod/openshift-ai-inference-openshift-default-8699b59bbc-4gqnw condition met [e2e-llm-inference-service] Pod(s) -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" are ready. [e2e-llm-inference-service] ⏳ Installing RHCL(Kuadrant) operator [e2e-llm-inference-service] namespace/kuadrant-system created [e2e-llm-inference-service] subscription.operators.coreos.com/rhcl-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/kuadrant created [e2e-llm-inference-service] Waiting for rhcl-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (0/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (5/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (10/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (15/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.2 found, but not yet Succeeded (Phase: Pending). Waiting... (20/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.2 found, but not yet Succeeded (Phase: Installing). Waiting... (25/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.2 found, but not yet Succeeded (Phase: Installing). Waiting... (30/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.2 found, but not yet Succeeded (Phase: Installing). Waiting... (35/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.2 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD kuadrants.kuadrant.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD kuadrants.kuadrant.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/kuadrants.kuadrant.io condition met [e2e-llm-inference-service] Waiting for apiserver discovery /apis/kuadrant.io/v1beta1 to list kuadrants (timeout: 120s)… [e2e-llm-inference-service] Discovery for kuadrant.io/v1beta1 includes kuadrants. [e2e-llm-inference-service] ⏳ sleeping 30s after discovery (RESTMapper can trail discovery)… [e2e-llm-inference-service] kuadrant.kuadrant.io/kuadrant created [e2e-llm-inference-service] ⏳ waiting for Kuadrant Ready (attempt 1/2, timeout 5m)… [e2e-llm-inference-service] kuadrant.kuadrant.io/kuadrant condition met [e2e-llm-inference-service] Waiting for pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" to be created... [e2e-llm-inference-service] Pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" found. [e2e-llm-inference-service] Current pods for -l "control-plane=authorino-operator" in namespace "kuadrant-system": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] authorino-operator-56fb6f6856-8wlrf 1/1 Running 0 77s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" to become ready... [e2e-llm-inference-service] pod/authorino-operator-56fb6f6856-8wlrf condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" are ready. [e2e-llm-inference-service] ⏳ waiting for authorino service to be created... [e2e-llm-inference-service] service/authorino-authorino-authorization condition met [e2e-llm-inference-service] service/authorino-authorino-authorization annotated [e2e-llm-inference-service] Warning: resource authorinos/authorino is missing the kubectl.kubernetes.io/last-applied-configuration annotation which is required by oc apply. oc apply should only be used on resources created declaratively by either oc create --save-config or oc apply. The missing annotation will be patched automatically. [e2e-llm-inference-service] authorino.operator.authorino.kuadrant.io/authorino configured [e2e-llm-inference-service] Waiting for pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" to be created... [e2e-llm-inference-service] Pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" found. [e2e-llm-inference-service] Current pods for -l "control-plane=authorino-operator" in namespace "kuadrant-system": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] authorino-operator-56fb6f6856-8wlrf 1/1 Running 0 86s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" to become ready... [e2e-llm-inference-service] pod/authorino-operator-56fb6f6856-8wlrf condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" are ready. [e2e-llm-inference-service] ✅ kuadrant(authorino) installed [e2e-llm-inference-service] Setting up Jaeger (Helm) for LLMISVC tracing e2e... [e2e-llm-inference-service] ⏳ Installing Jaeger All-in-One (Helm) into namespace observability [e2e-llm-inference-service] [INFO] Installing Helm v3.16.3 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed Helm v3.16.3 to /workspace/source/bin/helm [e2e-llm-inference-service] WARNING: Kubernetes configuration file is group-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] WARNING: Kubernetes configuration file is world-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] version.BuildInfo{Version:"v3.16.3", GitCommit:"cfd07493f46efc9debd9cc1b02a0961186df7fdf", GitTreeState:"clean", GoVersion:"go1.22.7"} [e2e-llm-inference-service] [INFO] Adding Jaeger Helm repository... [e2e-llm-inference-service] WARNING: Kubernetes configuration file is group-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] WARNING: Kubernetes configuration file is world-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] "jaegertracing" has been added to your repositories [e2e-llm-inference-service] [INFO] Installing Jaeger All-in-One 4.7.0... [e2e-llm-inference-service] WARNING: Kubernetes configuration file is group-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] WARNING: Kubernetes configuration file is world-readable. This is insecure. Location: /credentials/cluster-kh8hc-kubeconfig [e2e-llm-inference-service] NAME: jaeger [e2e-llm-inference-service] LAST DEPLOYED: Tue Jul 28 16:19:09 2026 [e2e-llm-inference-service] NAMESPACE: observability [e2e-llm-inference-service] STATUS: deployed [e2e-llm-inference-service] REVISION: 1 [e2e-llm-inference-service] TEST SUITE: None [e2e-llm-inference-service] NOTES: [e2e-llm-inference-service] ################################################################### [e2e-llm-inference-service] ### ⚠️ EXPERIMENTAL - NO STABILITY GUARANTEES ### [e2e-llm-inference-service] ### ### [e2e-llm-inference-service] ### This chart is under active development. ### [e2e-llm-inference-service] ### Breaking changes may occur in minor versions. ### [e2e-llm-inference-service] ### ### [e2e-llm-inference-service] ### See README.md for configuration details. ### [e2e-llm-inference-service] ################################################################### [e2e-llm-inference-service] [e2e-llm-inference-service] 🚀 Congratulations on successfully installing Jaeger v2.17.0 (Chart v4.7.0)! [e2e-llm-inference-service] [e2e-llm-inference-service] To access the query UI: [e2e-llm-inference-service] export POD_NAME=$(kubectl get pods --namespace observability -l "app.kubernetes.io/instance=jaeger,app.kubernetes.io/component=all-in-one" -o jsonpath="{.items[0].metadata.name}") [e2e-llm-inference-service] kubectl port-forward --namespace observability $POD_NAME 16686:16686 --address 0.0.0.0 & [e2e-llm-inference-service] Visit http://127.0.0.1:16686/ [e2e-llm-inference-service] [SUCCESS] Successfully installed Jaeger All-in-One via Helm [e2e-llm-inference-service] [INFO] Waiting for pods with label 'app.kubernetes.io/name=jaeger' in namespace 'observability' to be created... [e2e-llm-inference-service] [INFO] Found 1 pod(s) with label 'app.kubernetes.io/name=jaeger' [e2e-llm-inference-service] [INFO] Waiting for pods with label 'app.kubernetes.io/name=jaeger' in namespace 'observability' to be ready... [e2e-llm-inference-service] pod/jaeger-6bbd7d4c9d-444vm condition met [e2e-llm-inference-service] [SUCCESS] Pods with label 'app.kubernetes.io/name=jaeger' in namespace 'observability' are ready! [e2e-llm-inference-service] [SUCCESS] Jaeger is ready! [e2e-llm-inference-service] ⏳ Verifying Jaeger Service ports (OTLP 4317, Query 16686)... [e2e-llm-inference-service] ✅ Jaeger (Helm) ready — OTLP http://jaeger.observability.svc.cluster.local:4317 [e2e-llm-inference-service] Patching ingress domain... [e2e-llm-inference-service] configmap/inferenceservice-config patched [e2e-llm-inference-service] pod "kserve-controller-manager-66f44fdc75-htqsz" deleted [e2e-llm-inference-service] Waiting for kserve-controller-manager to be ready... [e2e-llm-inference-service] pod/kserve-controller-manager-66f44fdc75-mmkn2 condition met [e2e-llm-inference-service] Installing ODH Model Controller manually... [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/accounts.nim.opendatahub.io created [e2e-llm-inference-service] serviceaccount/model-serving-api created [e2e-llm-inference-service] serviceaccount/odh-model-controller created [e2e-llm-inference-service] role.rbac.authorization.k8s.io/leader-election-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/account-editor-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/account-viewer-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-prometheus-k8s created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/metrics-reader created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/model-serving-api created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/odh-model-controller-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/proxy-role created [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/leader-election-rolebinding created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/model-serving-api created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/odh-model-controller-rolebinding-opendatahub created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/proxy-rolebinding created [e2e-llm-inference-service] configmap/odh-model-controller-parameters created [e2e-llm-inference-service] service/model-serving-api created [e2e-llm-inference-service] service/odh-model-controller-metrics-service created [e2e-llm-inference-service] service/odh-model-controller-webhook-service created [e2e-llm-inference-service] deployment.apps/model-serving-api created [e2e-llm-inference-service] deployment.apps/odh-model-controller created [e2e-llm-inference-service] servicemonitor.monitoring.coreos.com/model-serving-api-metrics created [e2e-llm-inference-service] servicemonitor.monitoring.coreos.com/odh-model-controller-metrics-monitor created [e2e-llm-inference-service] template.template.openshift.io/autogluon-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/guardrails-detector-huggingface-serving-template created [e2e-llm-inference-service] template.template.openshift.io/kserve-ovms created [e2e-llm-inference-service] template.template.openshift.io/mlserver-cuda-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/mlserver-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template-fast-2 created [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/mutating.odh-model-controller.opendatahub.io created [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/validating.odh-model-controller.opendatahub.io created [e2e-llm-inference-service] Waiting for deployment "odh-model-controller" rollout to finish: 0 of 1 updated replicas are available... [e2e-llm-inference-service] deployment "odh-model-controller" successfully rolled out [e2e-llm-inference-service] networkpolicy.networking.k8s.io/allow-all created [e2e-llm-inference-service] KServe setup complete (namespace: kserve) [e2e-llm-inference-service] Add testing models to SeaweedFS S3 storage ... [e2e-llm-inference-service] Waiting for SeaweedFS deployment to be ready... [e2e-llm-inference-service] deployment "seaweedfs" successfully rolled out [e2e-llm-inference-service] S3 init job not completed, re-creating... [e2e-llm-inference-service] job.batch/s3-init replaced [e2e-llm-inference-service] Waiting for S3 init job to complete... [e2e-llm-inference-service] job.batch/s3-init condition met [e2e-llm-inference-service] Prepare CI namespace and install ServingRuntimes [e2e-llm-inference-service] Setting up CI namespace: kserve-ci-e2e-test [e2e-llm-inference-service] Tearing down CI namespace: kserve-ci-e2e-test [e2e-llm-inference-service] Namespace kserve-ci-e2e-test does not exist, skipping deletion [e2e-llm-inference-service] CI namespace teardown complete [e2e-llm-inference-service] Creating namespace kserve-ci-e2e-test [e2e-llm-inference-service] namespace/kserve-ci-e2e-test created [e2e-llm-inference-service] Applying S3 artifact secret [e2e-llm-inference-service] secret/mlpipeline-s3-artifact created [e2e-llm-inference-service] Applying storage-config secret [e2e-llm-inference-service] secret/storage-config created [e2e-llm-inference-service] Applying SeaweedFS S3 credentials secret [e2e-llm-inference-service] secret/seaweedfs-s3-creds created [e2e-llm-inference-service] Linking seaweedfs-s3-creds to default service account [e2e-llm-inference-service] Creating odh-trusted-ca-bundle configmap [e2e-llm-inference-service] configmap/odh-trusted-ca-bundle created [e2e-llm-inference-service] Installing ServingRuntimes [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-autogluonserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-huggingfaceserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-huggingfaceserver-multinode created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-lgbserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-mlserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-paddleserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-pmmlserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-predictiveserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-sklearnserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-tensorflow-serving created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-torchserve created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-tritonserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-vllmserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-xgbserver created [e2e-llm-inference-service] CI namespace setup complete [e2e-llm-inference-service] Setup complete [e2e-llm-inference-service] === E2E cluster / operator summary === [e2e-llm-inference-service] Client Version: 4.20.11 [e2e-llm-inference-service] Kustomize Version: v5.6.0 [e2e-llm-inference-service] Server Version: 4.21.26 [e2e-llm-inference-service] Kubernetes Version: v1.34.9 [e2e-llm-inference-service] ClusterVersion desired: 4.21.26 [e2e-llm-inference-service] ClusterVersion history (latest): 4.21.26 (Completed) [e2e-llm-inference-service] CSVs in kuadrant-system: [e2e-llm-inference-service] authorino-operator.v1.4.2 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.20.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.1 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.1 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.2 Succeeded [e2e-llm-inference-service] CSVs in openshift-keda: [e2e-llm-inference-service] authorino-operator.v1.4.2 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.20.0 Succeeded [e2e-llm-inference-service] custom-metrics-autoscaler.v2.19.0-2 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.1 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.1 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.2 Succeeded [e2e-llm-inference-service] CSVs in cert-manager-operator: [e2e-llm-inference-service] authorino-operator.v1.4.2 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.20.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.1 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.1 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.2 Succeeded [e2e-llm-inference-service] CSVs in openshift-lws-operator: [e2e-llm-inference-service] authorino-operator.v1.4.2 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.20.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.1 Succeeded [e2e-llm-inference-service] leader-worker-set.v1.0.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.1 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.2 Succeeded [e2e-llm-inference-service] CSVs in openshift-operators (ODH / shared operators, filtered): [e2e-llm-inference-service] authorino-operator.v1.4.2 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.1 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.1 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.2 Succeeded [e2e-llm-inference-service] Kuadrant / Authorino (diagnostics): [e2e-llm-inference-service] CRD kuadrants.kuadrant.io versions: v1beta1 served=true storage=true [e2e-llm-inference-service] Subscriptions in kuadrant-system: [e2e-llm-inference-service] authorino-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators authorino-operator.v1.4.2 [e2e-llm-inference-service] dns-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators dns-operator.v1.4.1 [e2e-llm-inference-service] limitador-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators limitador-operator.v1.4.1 [e2e-llm-inference-service] rhcl-operator stable redhat-operators rhcl-operator.v1.4.2 [e2e-llm-inference-service] Kuadrant CR conditions (kuadrant/kuadrant-system): [e2e-llm-inference-service] Ready=True (Ready) [e2e-llm-inference-service] KServe deployments in kserve: [e2e-llm-inference-service] kserve-controller-manager: ready=1 image=quay.io/opendatahub/kserve-controller@sha256:a627ec30debc454af106c2c34af695e96eccb25d564ea70da2ca083e0bc41d91 [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-controller@sha256:406c6b521a73de6d8758208f8b01fbda1aca8cbb849ba76749aa41667d674b0d [e2e-llm-inference-service] odh-model-controller: ready=1 image=quay.io/opendatahub/odh-model-controller:odh-incubating [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-model-controller@sha256:b5c835d640b650f164ba7f36d7008760c88e2974efa65853ee167323c441580c [e2e-llm-inference-service] llmisvc-controller-manager: ready=1 image=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:c5b450eea4db6d5fde1b2fb72cb1a1fa1929d109024104cb2f0ea7314a0fdd04 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:026af0bb86b6cacaeb05cf3d0aa19a498e6cede4c4d2792c417b5c2a7b81d23f [e2e-llm-inference-service] === End E2E cluster / operator summary === [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] CA certificate extracted [e2e-llm-inference-service] REQUESTS_CA_BUNDLE=/tmp/ca.crt [e2e-llm-inference-service] Run E2E tests: llminferenceservice and cluster_cpu and not autoscaling and not tracing [e2e-llm-inference-service] Starting E2E functional tests ... [e2e-llm-inference-service] Parallelism requested for pytest is 2 [e2e-llm-inference-service] ============================= test session starts ============================== [e2e-llm-inference-service] platform linux -- Python 3.11.13, pytest-7.4.4, pluggy-1.5.0 -- /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] cachedir: .pytest_cache [e2e-llm-inference-service] metadata: {'Python': '3.11.13', 'Platform': 'Linux-5.14.0-570.124.1.el9_6.x86_64-x86_64-with-glibc2.34', 'Packages': {'pytest': '7.4.4', 'pluggy': '1.5.0'}, 'Plugins': {'asyncio': '0.23.8', 'anyio': '4.9.0', 'xdist': '3.6.1', 'httpx': '0.30.0', 'cov': '5.0.0', 'json-report': '1.5.0', 'metadata': '3.1.1'}, 'PLATFORM': 'el9'} [e2e-llm-inference-service] rootdir: /workspace/source/test/e2e [e2e-llm-inference-service] configfile: pytest.ini [e2e-llm-inference-service] plugins: asyncio-0.23.8, anyio-4.9.0, xdist-3.6.1, httpx-0.30.0, cov-5.0.0, json-report-1.5.0, metadata-3.1.1 [e2e-llm-inference-service] asyncio: mode=Mode.STRICT [e2e-llm-inference-service] created: 2/2 workers [e2e-llm-inference-service] 2 workers [73 items] [e2e-llm-inference-service] [e2e-llm-inference-service] scheduling tests via WorkStealingScheduling [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-utilization-detector] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] 2026-07-28 16:21:39.130 6513 kserve INFO [conftest.py:configure_logger():40] Logger configured [e2e-llm-inference-service] 2026-07-28 16:21:39.131 6516 kserve INFO [conftest.py:configure_logger():40] Logger configured [e2e-llm-inference-service] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-utilization-detector] [e2e-llm-inference-service] llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-concurrency-detector] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-concurrency-detector] [e2e-llm-inference-service] llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-with-section-name] 2026-07-28 16:24:34.962 6513 kserve.trace Checking Gateway router-gateway-1 in namespace e2e-test-gateway-section-name-propagation-49b5edd9 [e2e-llm-inference-service] 2026-07-28 16:24:34.962 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace e2e-test-gateway-section-name-propagation-49b5edd9 [e2e-llm-inference-service] 2026-07-28 16:24:34.982 6513 kserve.trace Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:34.982 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():62] Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:34.991 6513 kserve.trace ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:34.991 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():70] ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-with-section-name] [e2e-llm-inference-service] llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-without-section-name] 2026-07-28 16:24:50.788 6513 kserve.trace Checking Gateway router-gateway-1 in namespace e2e-test-gateway-section-name-propagation-94799d44 [e2e-llm-inference-service] 2026-07-28 16:24:50.788 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace e2e-test-gateway-section-name-propagation-94799d44 [e2e-llm-inference-service] 2026-07-28 16:24:50.811 6513 kserve.trace Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:50.811 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():62] Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:50.822 6513 kserve.trace ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-28 16:24:50.822 6513 kserve.trace INFO [gw_api.py:create_or_update_gateway():70] ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-without-section-name] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_enabled_requires_token[cluster_cpu-cluster_single_node-auth-enabled-default] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_enabled_requires_token[cluster_cpu-cluster_single_node-auth-enabled-default] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_invalid_token_rejected[cluster_cpu-cluster_single_node-auth-invalid-token] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_invalid_token_rejected[cluster_cpu-cluster_single_node-auth-invalid-token] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_disabled_no_token_required[cluster_cpu-cluster_single_node-auth-disabled] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_disabled_no_token_required[cluster_cpu-cluster_single_node-auth-disabled] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-pd-config-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-pd-config-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-nonzero-threshold-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-nonzero-threshold-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-tokenizer-kvcache-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-tokenizer-kvcache-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_config_finalizer_added [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_config_finalizer_added [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_blocked_when_referenced [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_blocked_when_referenced [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_allowed_when_unreferenced [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_allowed_when_unreferenced [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_unblocked_after_service_deleted [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_config_deletion_unblocked_after_service_deleted [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_well_known_config_deletion_prevented_by_webhook [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_well_known_config_deletion_prevented_by_webhook [e2e-llm-inference-service] llmisvc/test_llm_inference_service_config_deletion.py::test_well_known_config_deletion_blocked_by_implicit_reference [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_config_deletion.py::test_well_known_config_deletion_blocked_by_implicit_reference [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha1_to_v1alpha2_conversion [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha1_to_v1alpha2_conversion [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha2_to_v1alpha1_conversion [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha2_to_v1alpha1_conversion [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_criticality_preservation_via_annotations [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_criticality_preservation_via_annotations [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_lora_criticality_preservation [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_lora_criticality_preservation [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_round_trip_conversion_preserves_fields [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_round_trip_conversion_preserves_fields [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_cleanup_hpa[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-single-lora-adapter-hf] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-single-lora-adapter-hf] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-multiple-lora-adapters] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_cleanup_hpa[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-multiple-lora-adapters] [e2e-llm-inference-service] llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] ERROR llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [e2e-llm-inference-service] ==================================== ERRORS ==================================== [e2e-llm-inference-service] _ ERROR at setup of test_prestop_hook[router-managed-workload-single-cpu-model-fb-opt-125m] _ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] llm_config = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_or_update_llmisvc_config(kserve_client, llm_config, namespace=None): [e2e-llm-inference-service] """Create or update an LLMInferenceServiceConfig resource.""" [e2e-llm-inference-service] version = llm_config["apiVersion"].split("/")[1] [e2e-llm-inference-service] [e2e-llm-inference-service] if namespace is None: [e2e-llm-inference-service] namespace = llm_config.get("metadata", {}).get("namespace", "default") [e2e-llm-inference-service] [e2e-llm-inference-service] name = llm_config.get("metadata", {}).get("name") [e2e-llm-inference-service] if not name: [e2e-llm-inference-service] raise ValueError("LLMInferenceServiceConfig must have a name in metadata") [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Checking LLMInferenceServiceConfig {name} in namespace {namespace}") [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > existing_config = kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1721: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1632: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-prestop-hook-tes-667dca99', 'namespace': 'e2e-test-prestop-hook-22727e61', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'name' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method get_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'name' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('name' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['name'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `name` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] if 'name' in local_var_params: [e2e-llm-inference-service] path_params['name'] = local_var_params['name'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1739: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-prestop-hook-tes-667dca99', 'namespace': 'e2e-test-prestop-hook-22727e61', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs/router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'e2e-test-prestop-hook-22727e61'), ('plural', 'llminferenceserviceconfigs'), ('name', 'router-managed-prestop-hook-tes-667dca99')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs/router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs/router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs/router-managed-prestop-hook-tes-667dca99' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (404) [e2e-llm-inference-service] E Reason: Not Found [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': '078b7c7c-4b4c-4af2-8959-e34b1a94ebd6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:08 GMT', 'Content-Length': '338'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"llminferenceserviceconfigs.serving.kserve.io \"router-managed-prestop-hook-tes-667dca99\" not found","reason":"NotFound","details":{"name":"router-managed-prestop-hook-tes-667dca99","group":"serving.kserve.io","kind":"llminferenceserviceconfigs"},"code":404} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] request = > [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.fixture(autouse=True) [e2e-llm-inference-service] def ensure_gateway_proxy_memory(request): [e2e-llm-inference-service] """After test setup creates gateways, patch them for proxy memory.""" [e2e-llm-inference-service] if not GATEWAY_PROXY_MEMORY: [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # Let test_case (llmisvc) create gateways first [e2e-llm-inference-service] [e2e-llm-inference-service] if "test_case" in request.fixturenames: [e2e-llm-inference-service] > request.getfixturevalue("test_case") [e2e-llm-inference-service] [e2e-llm-inference-service] common/gateway_proxy_istio.py:183: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] request = > [e2e-llm-inference-service] test_namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.fixture(scope="function") [e2e-llm-inference-service] def test_case(request, test_namespace): [e2e-llm-inference-service] tc = request.param [e2e-llm-inference-service] ns = test_namespace [e2e-llm-inference-service] [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] tc.namespace = ns [e2e-llm-inference-service] for peer in tc.peers: [e2e-llm-inference-service] peer.namespace = ns [e2e-llm-inference-service] [e2e-llm-inference-service] for func in tc.before_test: [e2e-llm-inference-service] func(tc) [e2e-llm-inference-service] [e2e-llm-inference-service] > _setup_test_case_service(kserve_client, tc, request.node.name, namespace=ns) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1611: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] tc = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=...test-prestop-hook-22727e61', before_test=[], after_test=[], peers=[], llm_service=None, model_name='facebook/opt-125m') [e2e-llm-inference-service] test_node_name = 'test_prestop_hook[router-managed-workload-single-cpu-model-fb-opt-125m]' [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61', peer_index = None [e2e-llm-inference-service] [e2e-llm-inference-service] def _setup_test_case_service( [e2e-llm-inference-service] kserve_client, tc, test_node_name, namespace, peer_index=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Create LLMInferenceServiceConfigs and build the LLMInferenceService for a TestCase. [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a list of created config names for cleanup tracking. [e2e-llm-inference-service] """ [e2e-llm-inference-service] missing_refs = [ [e2e-llm-inference-service] ref for ref in tc.base_refs if ref not in LLMINFERENCESERVICE_CONFIGS [e2e-llm-inference-service] ] [e2e-llm-inference-service] if missing_refs: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Missing base_refs in LLMINFERENCESERVICE_CONFIGS: {missing_refs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] if not tc.service_name: [e2e-llm-inference-service] suffix = f"-peer-{peer_index}" if peer_index is not None else "" [e2e-llm-inference-service] tc.service_name = generate_service_name(test_node_name + suffix, tc.base_refs) [e2e-llm-inference-service] if tc.model_name == "default/model": [e2e-llm-inference-service] tc.model_name = _get_model_name_from_configs(tc.base_refs) [e2e-llm-inference-service] elif "{namespace}" in tc.model_name: [e2e-llm-inference-service] tc.model_name = tc.model_name.format(namespace=namespace) [e2e-llm-inference-service] [e2e-llm-inference-service] created_configs = [] [e2e-llm-inference-service] unique_base_refs = [] [e2e-llm-inference-service] for base_ref in tc.base_refs: [e2e-llm-inference-service] unique_config_name = generate_k8s_safe_suffix(base_ref, [tc.service_name]) [e2e-llm-inference-service] unique_base_refs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] config = LLMINFERENCESERVICE_CONFIGS[base_ref] [e2e-llm-inference-service] spec = config(namespace) if callable(config) else copy.deepcopy(config) [e2e-llm-inference-service] [e2e-llm-inference-service] unique_config_body = { [e2e-llm-inference-service] "apiVersion": "serving.kserve.io/v1alpha1", [e2e-llm-inference-service] "kind": "LLMInferenceServiceConfig", [e2e-llm-inference-service] "metadata": { [e2e-llm-inference-service] "name": unique_config_name, [e2e-llm-inference-service] "namespace": namespace, [e2e-llm-inference-service] }, [e2e-llm-inference-service] "spec": spec, [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] > _create_or_update_llmisvc_config(kserve_client, unique_config_body, namespace) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1573: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] llm_config = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_or_update_llmisvc_config(kserve_client, llm_config, namespace=None): [e2e-llm-inference-service] """Create or update an LLMInferenceServiceConfig resource.""" [e2e-llm-inference-service] version = llm_config["apiVersion"].split("/")[1] [e2e-llm-inference-service] [e2e-llm-inference-service] if namespace is None: [e2e-llm-inference-service] namespace = llm_config.get("metadata", {}).get("namespace", "default") [e2e-llm-inference-service] [e2e-llm-inference-service] name = llm_config.get("metadata", {}).get("name") [e2e-llm-inference-service] if not name: [e2e-llm-inference-service] raise ValueError("LLMInferenceServiceConfig must have a name in metadata") [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Checking LLMInferenceServiceConfig {name} in namespace {namespace}") [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] existing_config = kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llm_config["metadata"] = existing_config["metadata"] [e2e-llm-inference-service] [e2e-llm-inference-service] outputs = kserve_client.api_instance.replace_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] llm_config, [e2e-llm-inference-service] ) [e2e-llm-inference-service] logger.info(f"✓ Successfully updated LLMInferenceServiceConfig {name}") [e2e-llm-inference-service] return outputs [e2e-llm-inference-service] [e2e-llm-inference-service] except client.rest.ApiException as e: [e2e-llm-inference-service] if e.status == 404: # Not found - create it [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"Resource not found, creating LLMInferenceServiceConfig {name}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] > outputs = kserve_client.api_instance.create_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] llm_config, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1747: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def create_namespaced_custom_object(self, group, version, namespace, plural, body, **kwargs): # noqa: E501 [e2e-llm-inference-service] """create_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Creates a namespace scoped Custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.create_namespaced_custom_object(group, version, namespace, plural, body, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: The custom resource's group name (required) [e2e-llm-inference-service] :param str version: The custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: The custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param object body: The JSON schema of the Resource to create. (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. [e2e-llm-inference-service] :param str dry_run: When present, indicates that modifications should not be persisted. An invalid or unrecognized dryRun directive will result in an error response and no further processing of the request. Valid values are: - All: all dry run stages will be processed [e2e-llm-inference-service] :param str field_manager: fieldManager is a name associated with the actor or entity that is making these changes. The value must be less than or 128 characters long, and only contain printable characters, as defined by https://golang.org/pkg/unicode/#IsPrint. [e2e-llm-inference-service] :param str field_validation: fieldValidation instructs the server on how to handle objects in the request (POST/PUT/PATCH) containing unknown or duplicate fields. Valid values are: - Ignore: This will ignore any unknown fields that are silently dropped from the object, and will ignore all but the last duplicate field that the decoder encounters. This is the default behavior prior to v1.23. - Warn: This will send a warning via the standard warning response header for each unknown field that is dropped from the object, and for each duplicate field that is encountered. The request will still succeed if there are no other errors, and will only persist the last of any duplicate fields. This is the default in v1.23+ - Strict: This will fail the request with a BadRequest error if any unknown fields would be dropped from the object, or if any duplicate fields are present. The error returned from the server will contain all unknown and duplicate fields encountered. (optional) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.create_namespaced_custom_object_with_http_info(group, version, namespace, plural, body, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:231: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-prestop-hook-22727e61' [e2e-llm-inference-service] plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'body', 'pretty', ...], 'au...k-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [...]}}}}}, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'body', 'pretty', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'namespace': 'e2e-test-prestop-hook-22727e61', 'plural': 'llminferenceserviceconfigs', 'version': 'v1alpha1'} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def create_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, body, **kwargs): # noqa: E501 [e2e-llm-inference-service] """create_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Creates a namespace scoped Custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.create_namespaced_custom_object_with_http_info(group, version, namespace, plural, body, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: The custom resource's group name (required) [e2e-llm-inference-service] :param str version: The custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: The custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param object body: The JSON schema of the Resource to create. (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. [e2e-llm-inference-service] :param str dry_run: When present, indicates that modifications should not be persisted. An invalid or unrecognized dryRun directive will result in an error response and no further processing of the request. Valid values are: - All: all dry run stages will be processed [e2e-llm-inference-service] :param str field_manager: fieldManager is a name associated with the actor or entity that is making these changes. The value must be less than or 128 characters long, and only contain printable characters, as defined by https://golang.org/pkg/unicode/#IsPrint. [e2e-llm-inference-service] :param str field_validation: fieldValidation instructs the server on how to handle objects in the request (POST/PUT/PATCH) containing unknown or duplicate fields. Valid values are: - Ignore: This will ignore any unknown fields that are silently dropped from the object, and will ignore all but the last duplicate field that the decoder encounters. This is the default behavior prior to v1.23. - Warn: This will send a warning via the standard warning response header for each unknown field that is dropped from the object, and for each duplicate field that is encountered. The request will still succeed if there are no other errors, and will only persist the last of any duplicate fields. This is the default in v1.23+ - Strict: This will fail the request with a BadRequest error if any unknown fields would be dropped from the object, or if any duplicate fields are present. The error returned from the server will contain all unknown and duplicate fields encountered. (optional) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'body', [e2e-llm-inference-service] 'pretty', [e2e-llm-inference-service] 'dry_run', [e2e-llm-inference-service] 'field_manager', [e2e-llm-inference-service] 'field_validation' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method create_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'body' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('body' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['body'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `body` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] if 'pretty' in local_var_params and local_var_params['pretty'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('pretty', local_var_params['pretty'])) # noqa: E501 [e2e-llm-inference-service] if 'dry_run' in local_var_params and local_var_params['dry_run'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('dryRun', local_var_params['dry_run'])) # noqa: E501 [e2e-llm-inference-service] if 'field_manager' in local_var_params and local_var_params['field_manager'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('fieldManager', local_var_params['field_manager'])) # noqa: E501 [e2e-llm-inference-service] if 'field_validation' in local_var_params and local_var_params['field_validation'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('fieldValidation', local_var_params['field_validation'])) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] if 'body' in local_var_params: [e2e-llm-inference-service] body_params = local_var_params['body'] [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}', 'POST', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:354: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}' [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'namespace': 'e2e-test-prestop-hook-22727e61', 'plural': 'llminferenceserviceconfigs', 'version': 'v1alpha1'} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs' [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'e2e-test-prestop-hook-22727e61'), ('plural', 'llminferenceserviceconfigs')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [] [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif method == "HEAD": [e2e-llm-inference-service] return self.rest_client.HEAD(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif method == "OPTIONS": [e2e-llm-inference-service] return self.rest_client.OPTIONS(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] elif method == "POST": [e2e-llm-inference-service] > return self.rest_client.POST(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] post_params=post_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:391: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], post_params = [] [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def POST(self, url, headers=None, query_params=None, post_params=None, [e2e-llm-inference-service] body=None, _preload_content=True, _request_timeout=None): [e2e-llm-inference-service] > return self.request("POST", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] post_params=post_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] body=body) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:279: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-prestop-hook-22727e61/llminferenceserviceconfigs' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...-hook-22727e61'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {'template': {'containers': [{...}]}}}}} [e2e-llm-inference-service] post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': 'ef750a44-ed5d-408a-9cf6-fbde79e2df1e', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:08 GMT', 'Content-Length': '701'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Internal error occurred: failed calling webhook \"llminferenceserviceconfig.kserve-webhook-server.v1alpha1.validator\": failed to call webhook: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/validate-serving-kserve-io-v1alpha1-llminferenceserviceconfig?timeout=10s\": EOF","reason":"InternalError","details":{"causes":[{"message":"failed calling webhook \"llminferenceserviceconfig.kserve-webhook-server.v1alpha1.validator\": failed to call webhook: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/validate-serving-kserve-io-v1alpha1-llminferenceserviceconfig?timeout=10s\": EOF"}]},"code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-prestop-hook-22727e61 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-kserve-custom-ca-bundle from kserve-ci-e2e-test to e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-prestop-hook-tes-667dca99 in namespace e2e-test-prestop-hook-22727e61 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-prestop-hook-tes-667dca99 [e2e-llm-inference-service] =================================== FAILURES =================================== [e2e-llm-inference-service] _ test_llm_autoscaling_hpa_deployment[router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prom... {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.autoscaling_hpa [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator-no-replicas", [e2e-llm-inference-service] "prometheus-scrape", [e2e-llm-inference-service] "scaling-hpa", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="autoscale-hpa-deploy", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_autoscaling_hpa_deployment(test_case: TestCase): [e2e-llm-inference-service] """HPA + Deployment: HPA exists with WVA annotations; pods scale up under load.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = _new_kserve_client() [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] ns = test_case.namespace [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _create_and_wait(kserve_client, test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:542: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prom... {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_and_wait(kserve_client, test_case): [e2e-llm-inference-service] """Create LLMISVC and wait for it to be ready.""" [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:482: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...e-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T16:34:40.984987', start_time = 1785256480.9853034 [e2e-llm-inference-service] duration = 900.6659591197968, timestamp_end = '2026-07-28T16:49:41.651274' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....toscale-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f1be0342b60> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-autoscale-hpa-de-ec1dce8b in namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-autoscale-hpa-de-ec1dce8b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-autoscale-hpa-de-ec1dce8b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-llmd-simulator-no-repl-38916baa in namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-no-repl-38916baa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-no-repl-38916baa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-4c186bcf in namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-4c186bcf [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-4c186bcf [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-deplo-347a3180 in namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-deplo-347a3180 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-deplo-347a3180 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_autoscaling_hpa_deployment] [2026-07-28T16:34:40.920746] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prompt='KServe is a', service_name='autoscale-hpa-deploy', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-de-ec1dce8b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-38916baa'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T16:34:40.933324] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-de-ec1dce8b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-38916baa'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T16:34:40.984766] end - ✅ in 0.051s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T16:34:40.984987] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-de-ec1dce8b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-38916baa'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/autoscale-hpa-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'Inference Pool e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/autoscale-hpa-deploy-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/autoscale-hpa-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/autoscale-hpa-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'HPA conditions not yet available', 'reason': 'HPAProgressing', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 3: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 4: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 5: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T16:49:41.651274] end - ❌ 900.666s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T16:49:41.651467] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-de-ec1dce8b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-38916baa'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-4c186bcf'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-deplo-347a3180'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 2 pod(s) for autoscale-hpa-deploy still terminating: ['autoscale-hpa-deploy-kserve-66746bd8b9-qvz5j', 'autoscale-hpa-deploy-kserve-router-scheduler-5bbdf87d58-5t9zx'] [e2e-llm-inference-service] assert not ['autoscale-hpa-deploy-kserve-66746bd8b9-qvz5j', 'autoscale-hpa-deploy-kserve-router-scheduler-5bbdf87d58-5t9zx'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 1 pod(s) for autoscale-hpa-deploy still terminating: ['autoscale-hpa-deploy-kserve-router-scheduler-5bbdf87d58-5t9zx'] [e2e-llm-inference-service] assert not ['autoscale-hpa-deploy-kserve-router-scheduler-5bbdf87d58-5t9zx'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-28T16:50:02.043487] end - ✅ in 20.391s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_autoscaling_hpa_deployment] [2026-07-28T16:50:02.043627] end - ❌ 921.123s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:00Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:35:32Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-deploy-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-autoscaling-hpa-deployment-9ba6f3f4 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_autoscaling_keda_deployment[router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-keda] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-keda'], pro... {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.autoscaling_keda [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator-no-replicas", [e2e-llm-inference-service] "prometheus-scrape", [e2e-llm-inference-service] "scaling-keda", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="autoscale-keda-deploy", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_autoscaling_keda_deployment(test_case: TestCase): [e2e-llm-inference-service] """KEDA + Deployment: ScaledObject exists with WVA annotations; no HPA; pods scale up under load.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = _new_kserve_client() [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] ns = test_case.namespace [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _create_and_wait(kserve_client, test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:606: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-keda'], pro... {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_and_wait(kserve_client, test_case): [e2e-llm-inference-service] """Create LLMISVC and wait for it to be ready.""" [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:482: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...e-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T16:50:02.775761', start_time = 1785257402.776031 [e2e-llm-inference-service] duration = 900.7376863956451, timestamp_end = '2026-07-28T17:05:03.513732' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....toscale-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f1be01318a0> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-autoscaling-keda-deployment-b2150d0b with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-autoscale-keda-d-27e06c40 in namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-autoscale-keda-d-27e06c40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-autoscale-keda-d-27e06c40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-llmd-simulator-no-repl-da49e827 in namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-no-repl-da49e827 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-no-repl-da49e827 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-101f2a9d in namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-101f2a9d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-101f2a9d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig scaling-keda-autoscale-keda-dep-1ac84077 in namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig scaling-keda-autoscale-keda-dep-1ac84077 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig scaling-keda-autoscale-keda-dep-1ac84077 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_autoscaling_keda_deployment] [2026-07-28T16:50:02.712187] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-keda'], prompt='KServe is a', service_name='autoscale-keda-deploy', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-autoscaling-keda-deployment-b2150d0b', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-deployment-b2150d0b', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-d-27e06c40'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-da49e827'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T16:50:02.725462] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-deployment-b2150d0b', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-d-27e06c40'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-da49e827'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T16:50:02.775627] end - ✅ in 0.050s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T16:50:02.775761] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-deployment-b2150d0b', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-d-27e06c40'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-da49e827'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-deployment-b2150d0b/autoscale-keda-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'Inference Pool e2e-test-llm-autoscaling-keda-deployment-b2150d0b/autoscale-keda-deploy-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-deployment-b2150d0b/autoscale-keda-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-deployment-b2150d0b/autoscale-keda-deploy-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'ScaledObject conditions not yet available', 'reason': 'ScaledObjectProgressing', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 6: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 7: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 8: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:05:03.513732] end - ❌ 900.738s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T17:05:03.513965] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-deploy', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-deployment-b2150d0b', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-d-27e06c40'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-da49e827'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-101f2a9d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-dep-1ac84077'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 2 pod(s) for autoscale-keda-deploy still terminating: ['autoscale-keda-deploy-kserve-77fc6d7cf9-rhsdn', 'autoscale-keda-deploy-kserve-router-scheduler-79d6bcd899-r275g'] [e2e-llm-inference-service] assert not ['autoscale-keda-deploy-kserve-77fc6d7cf9-rhsdn', 'autoscale-keda-deploy-kserve-router-scheduler-79d6bcd899-r275g'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 1 pod(s) for autoscale-keda-deploy still terminating: ['autoscale-keda-deploy-kserve-router-scheduler-79d6bcd899-r275g'] [e2e-llm-inference-service] assert not ['autoscale-keda-deploy-kserve-router-scheduler-79d6bcd899-r275g'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-28T17:05:33.782552] end - ✅ in 30.268s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_autoscaling_keda_deployment] [2026-07-28T17:05:33.782702] end - ❌ 931.070s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:09Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T16:50:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:50:24Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-autoscaling-keda-deployment-b2150d0b (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _____ test_llm_inference_service[router-managed-workload-pd-cpu-model-pvc] _____ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-pd-cpu', 'model-pvc'], prompt='KServe is a', service_name='llmisvc-mod... {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway="router-gateway-1", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway="router-gateway-1", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] routes=[ [e2e-llm-inference-service] make_router_main_route( [e2e-llm-inference-service] "router-route-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] "router-with-refs-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] make_router_health_route( [e2e-llm-inference-service] "router-route-2", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] "router-with-refs-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway="router-gateway-2", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] routes=[ [e2e-llm-inference-service] make_router_main_route( [e2e-llm-inference-service] "router-route-3", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] "router-with-refs-pd-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] make_router_health_route( [e2e-llm-inference-service] "router-route-4", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] "router-with-refs-pd-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_scheduler_configmap(namespace=tc.namespace) [e2e-llm-inference-service] ], [e2e-llm-inference-service] after_test=[ [e2e-llm-inference-service] lambda tc: delete_scheduler_configmap(namespace=tc.namespace) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Standalone tokenizer — clean path: token-producer in inline config [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-tokenizer-kvcache", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="tokenizer-clean-path-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Standalone tokenizer — migration path: legacy precise-prefix-cache-scorer [e2e-llm-inference-service] # triggers auto-provisioned tokenizer without explicit tokenizer:{} field [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="tokenizer-migration-path-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name="publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "publishers/{namespace}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] "publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:866: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...del-p-9d807ba3'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T16:54:51.106256', start_time = 1785257691.1065323 [e2e-llm-inference-service] duration = 900.4667983055115, timestamp_end = '2026-07-28T17:09:51.573340' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....svc-model-p-9d807ba3'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f2fed4c4e00> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-inference-service-48639af5 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1959 Created PVC e2e-pvc-model-storage in namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:2099 Created model download Job e2e-pvc-model-download in namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:2124 Model download Job e2e-pvc-model-download completed successfully [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-llmisvc-model-pv-d968e7b0 in namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-llmisvc-model-pv-d968e7b0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-llmisvc-model-pv-d968e7b0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-pd-cpu-llmisvc-model-p-9d807ba3 in namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-pd-cpu-llmisvc-model-p-9d807ba3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-pd-cpu-llmisvc-model-p-9d807ba3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-49c1f027 in namespace e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-49c1f027 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-49c1f027 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-28T16:54:50.984800] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-pd-cpu', 'model-pvc'], prompt='KServe is a', service_name='llmisvc-model-pvc-router-manage-e8706282', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-inference-service-48639af5', before_test=[ at 0x7f2fedd960c0>], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-e8706282', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-48639af5', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-d968e7b0'}, [e2e-llm-inference-service] {'name': 'workload-pd-cpu-llmisvc-model-p-9d807ba3'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T16:54:50.997594] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-e8706282', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-48639af5', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-d968e7b0'}, [e2e-llm-inference-service] {'name': 'workload-pd-cpu-llmisvc-model-p-9d807ba3'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T16:54:51.106071] end - ✅ in 0.108s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T16:54:51.106256] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-e8706282', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-48639af5', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-d968e7b0'}, [e2e-llm-inference-service] {'name': 'workload-pd-cpu-llmisvc-model-p-9d807ba3'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-49c1f027'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:03Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'message': 'Inference Pool e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:09:51.573340] end - ❌ 900.467s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:903 [router-managed-workload-pd-cpu-model-pvc] ❌ ERROR: Failed to call llm inference service llmisvc-model-pvc-router-manage-e8706282: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:243 # Diagnostics for 'llmisvc-model-pvc-router-manage-e8706282' in 'e2e-test-llm-inference-service-48639af5' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:244 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:247 # LLMInferenceService llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:256 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:54:51Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:54:51Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:54:51Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler-latency-predictor: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tokenizer: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:prefill: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:57:20Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] resourceVersion: '60378' [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-llmisvc-model-pv-d968e7b0 [e2e-llm-inference-service] - name: workload-pd-cpu-llmisvc-model-p-9d807ba3 [e2e-llm-inference-service] - name: model-pvc-llmisvc-model-pvc-rou-49c1f027 [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler-latency-predictor: kserve-config-llm-scheduler-latency-predictor [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tokenizer: kserve-config-llm-tokenizer [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:12Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:12Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:57:20Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: MainWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:12Z' [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: PrefillWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:03Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:12Z' [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:40Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:40Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:12Z' [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:68 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:69 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-48639af5 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-48639af5/e2e-pvc-model-download-7hv22 to ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:36 e2e-test-llm-inference-service-48639af5 attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c81af45c-2687-4fb1-80d9-500c5cb5c636" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:37 e2e-test-llm-inference-service-48639af5 multus Normal AddedInterface Add eth0 [10.133.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:37 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:37 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:37 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:30 e2e-test-llm-inference-service-48639af5 job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-7hv22 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:49 e2e-test-llm-inference-service-48639af5 job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:30 e2e-test-llm-inference-service-48639af5 persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:30 e2e-test-llm-inference-service-48639af5 ebs.csi.aws.com_aws-ebs-csi-driver-controller-57f6947994-rxq7b_37543d0b-c2ce-44f2-9725-79e46f1d830a Normal Provisioning External provisioner is provisioning volume for claim "e2e-test-llm-inference-service-48639af5/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:30 e2e-test-llm-inference-service-48639af5 persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:54:32 e2e-test-llm-inference-service-48639af5 ebs.csi.aws.com_aws-ebs-csi-driver-controller-57f6947994-rxq7b_37543d0b-c2ce-44f2-9725-79e46f1d830a Normal ProvisioningSucceeded Successfully provisioned volume pvc-c81af45c-2687-4fb1-80d9-500c5cb5c636 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-48639af5 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m to ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:06 e2e-test-llm-inference-service-48639af5 attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c81af45c-2687-4fb1-80d9-500c5cb5c636" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 multus Normal AddedInterface Add eth0 [10.133.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Pulled Container image "quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Pulled Container image "vllm/vllm-openai-cpu:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:08 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:56:58 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-139-138.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.133.0.54:8001/health": dial tcp 10.133.0.54:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b8x7fs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-48639af5 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b8x7fs to ip-10-0-140-86.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:06 e2e-test-llm-inference-service-48639af5 attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c81af45c-2687-4fb1-80d9-500c5cb5c636" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b89cd7957 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.Secret e2e-test-llm-inference-service-48639af5/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.ServiceAccount e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.Role e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.RoleBinding e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.Deployment e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.Deployment e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.Service e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.ServiceAccount e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /e2e-test-llm-inference-service-132820f03254470cd8a0146008974a5f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:11 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:05:14 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Updated Updated v1.Secret e2e-test-llm-inference-service-48639af5/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:11 e2e-test-llm-inference-service-48639af5 LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-48639af5 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-48639af5/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 to ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:02 e2e-test-llm-inference-service-48639af5 multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:02 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-142-94.ec2.internal Normal Pulled Container image "quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:02 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-142-94.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:02 e2e-test-llm-inference-service-48639af5 kubelet/ip-10-0-142-94.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 16:55:01 e2e-test-llm-inference-service-48639af5 deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-7fc69df7b8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:172 ### Pod llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### init-container 'llm-d-routing-sidecar' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:225 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:226 {"level":"info","ts":1785257708.4293783,"msg":"Proxy starting","Built on":"","From Git SHA":"unknown"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.4294071,"msg":"Proxy configuration","config":"{\"Port\":\"8000\",\"KVConnector\":\"nixlv2\",\"ECConnector\":\"\",\"DataParallelSize\":1,\"MaxIdleConnsPerHost\":1024,\"EnablePrefillerSampling\":false,\"PrefillMaxRetries\":0,\"PrefillRetryBackoff\":200000000,\"UseTLSForPrefiller\":true,\"UseTLSForDecoder\":true,\"UseTLSForEncoder\":false,\"InsecureSkipVerifyForPrefiller\":false,\"InsecureSkipVerifyForEncoder\":false,\"InsecureSkipVerifyForDecoder\":false,\"SecureServing\":true,\"CertPath\":\"/var/run/kserve/tls\",\"MooncakeBootstrapPort\":8998,\"EnableSSRFProtection\":true,\"InferencePoolNamespace\":\"e2e-test-llm-inference-service-48639af5\",\"InferencePoolName\":\"llmisvc-model-pvc-router-manage-e8706282-inference-pool\",\"PoolGroup\":\"inference.networking.x-k8s.io\",\"DecodeChunkSize\":0,\"Tracing\":false,\"MoRIIOWriteMode\":false,\"MoRIIODecodeNotifyPort\":61005,\"MoRIIODecodeHandshakePort\":6301,\"MoRIIODecodePodIP\":\"\",\"MoRIIOParallelDispatch\":false,\"MoRIIOPrefillHandshakePort\":6301,\"MoRIIOPrefillNotifyPort\":61005,\"MoRIIOTPSize\":1,\"MoRIIODPSize\":1,\"MoRIIORemoteHosts\":null,\"MoRIIODPSizeLocal\":0,\"MoRIIODecodeHosts\":null,\"DecoderURL\":\"https://localhost:8001\"}"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.4300056,"logger":"allowlist-validator","msg":"starting SSRF protection allowlist validator","namespace":"e2e-test-llm-inference-service-48639af5","poolName":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","gvr":"inference.networking.x-k8s.io/v1alpha2, Resource=inferencepools"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.4507275,"logger":"allowlist-validator","msg":"InferencePool added","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.4678795,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":0,"targets":{}} [e2e-llm-inference-service] {"level":"info","ts":1785257708.467953,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":0,"targets":{}} [e2e-llm-inference-service] {"level":"info","ts":1785257708.5301585,"logger":"allowlist-validator","msg":"allowlist validator started successfully"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.5312514,"logger":"proxy server on port 8000","msg":"server TLS configured"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.5312662,"logger":"proxy server on port 8000","msg":"starting","addr":"[::]:8000"} [e2e-llm-inference-service] {"level":"info","ts":1785257708.7573879,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257709.7644596,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257718.2477682,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257738.4515724,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257738.462999,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257738.4630723,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257768.4524164,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257768.4653573,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257768.4654357,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257798.4524975,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257798.4675508,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257798.4676282,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257828.250075,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257828.2689695,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257828.452886,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257828.4639168,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257828.4639757,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257858.453731,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257858.4644556,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257858.4645073,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257888.4539676,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257888.4658418,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257888.465906,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257918.454819,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257918.465913,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257918.4659724,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257948.455881,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257948.4695964,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257948.4696527,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257978.4560604,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785257978.467154,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257978.4672086,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258008.4565554,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258008.469512,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258008.4695685,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258038.457245,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258038.4695964,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258038.469654,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258068.4581177,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258068.4706433,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258068.4707084,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258098.4585593,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258098.471416,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258098.471474,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258128.459043,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258128.4765851,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258128.476649,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258158.4596105,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258158.475259,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258158.4753199,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258188.4599178,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258188.4737606,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258188.4738407,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258218.4606128,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258218.4728642,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258218.4729207,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258248.4613929,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258248.4728584,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258248.4729125,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258278.4616683,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258278.4728987,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258278.472951,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258308.4624388,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258308.475135,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258308.4751968,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258338.4628985,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258338.4757164,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258338.4757755,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258368.4636817,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258368.4769986,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258368.4770539,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258398.464606,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258398.4772446,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258398.4773076,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258428.4655192,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258428.4780395,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258428.4780934,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258458.4662256,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258458.4784942,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258458.4785497,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258488.466325,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258488.4782255,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258488.478284,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258518.4664364,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258518.4780025,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258518.478059,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258548.4666653,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258548.47809,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258548.4781487,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258578.4667032,"logger":"allowlist-validator","msg":"InferencePool updated","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool"} [e2e-llm-inference-service] {"level":"info","ts":1785258578.478951,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785258578.4790072,"logger":"allowlist-validator","msg":"rebuilt allowlist","targetCount":2,"targets":{"10.133.0.54":{},"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m":{}}} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:225 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:226 (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:236] auto thread-binding list (id, physical core): [(4, 0), (5, 1), (6, 2), (7, 3)] [e2e-llm-inference-service] [W728 16:56:18.433049703 utils.cpp:76] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_threads_env) [e2e-llm-inference-service] [W728 16:56:18.433074828 utils.cpp:103] Warning: NUMA binding: Using MEMBIND policy for memory allocation on the NUMA nodes (0). Memory allocations will be strictly bound to these NUMA nodes. (function init_cpu_threads_env) [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] OMP threads binding of Process 74: [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] OMP tid: 74, core 4 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] OMP tid: 91, core 5 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] OMP tid: 92, core 6 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] OMP tid: 93, core 7 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [v1/worker/cpu_worker.py:109] [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:18 [distributed/parallel_state.py:1356] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.133.0.54:60197 backend=gloo [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [distributed/parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.133.0.54:60197 backend=gloo [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:18 [distributed/parallel_state.py:1459] Detected 1 nodes in the distributed environment [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:18 [distributed/parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [v1/sample/logits_processor/__init__.py:65] No logitsprocs plugins installed (group vllm.logits_processors). [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [model_executor/offloader/base.py:107] Offloader set to NoopOffloader (no offloading). [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:56:19 [v1/worker/cpu_model_runner.py:71] Starting to load model /mnt/models... [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [compilation/decorators.py:213] Inferred dynamic dimensions for forward method of : ['input_ids', 'positions', 'intermediate_tensors', 'inputs_embeds'] [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [config/compilation.py:1194] enabled custom ops: Counter() [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [config/compilation.py:1195] disabled custom ops: Counter({'vocab_parallel_embedding': 1, 'logits_processor': 1}) [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:56:19 [model_executor/model_loader/base_loader.py:63] Loading weights on cpu ... [e2e-llm-inference-service] (EngineCore pid=74) Loading pt checkpoint shards: 0% Completed | 0/1 [00:00 [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:56:28 [v1/engine/utils.py:1047] Waiting for 1 local, 0 remote core engine proc(s) to start. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:56:38 [v1/engine/utils.py:1047] Waiting for 1 local, 0 remote core engine proc(s) to start. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:56:48 [v1/engine/utils.py:1047] Waiting for 1 local, 0 remote core engine proc(s) to start. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:56:58 [v1/engine/utils.py:1047] Waiting for 1 local, 0 remote core engine proc(s) to start. [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:57:03 [compilation/decorators.py:640] saved AOT compiled function to /home/.cache/vllm/torch_compile_cache/torch_aot_compile/86c9c3c579382eef68a98ac1d59b39811ba08abef3b4e90675a32c8dec3d7c90/rank_0_0/model [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:57:04 [compilation/monitor.py:76] Initial profiling/warmup run took 1.17 s [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:57:04 [v1/worker/cpu_model_runner.py:92] Warming up done. [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:57:04 [v1/engine/core.py:283] init engine (profile, create kv cache, warmup model) took 44.17 seconds [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:57:04 [tokenizers/registry.py:68] Loading CachedHfTokenizer for tokenizer_mode='hf' [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:57:05 [utils/gc_utils.py:40] GC Debug Config. enabled:False,top_objects:-1 [e2e-llm-inference-service] (EngineCore pid=74) INFO 07-28 16:57:05 [config/vllm.py:790] Asynchronous scheduling is disabled. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [v1/engine/utils.py:1158] READY from local core engine process 0. [e2e-llm-inference-service] (EngineCore pid=74) WARNING 07-28 16:57:05 [config/vllm.py:859] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:57:05 [v1/engine/core.py:1158] EngineCore waiting for work. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [v1/metrics/loggers.py:273] Engine 000: vllm cache_config_info with initialization after num_gpu_blocks is: 227 [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:57:05 [v1/engine/core.py:1158] EngineCore waiting for work. [e2e-llm-inference-service] (EngineCore pid=74) DEBUG 07-28 16:57:05 [v1/engine/core.py:1158] EngineCore waiting for work. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/openai/api_server.py:590] Supported tasks: ['generate'] [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/base.py:197] Warming up chat template processing... [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] Failed to load AutoTokenizer chat template for /mnt/models [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] Traceback (most recent call last): [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] File "/opt/venv/lib/python3.12/site-packages/vllm/renderers/hf.py", line 120, in resolve_chat_template [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] return tokenizer.get_chat_template(chat_template, tools=tools) [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] File "/opt/venv/lib/python3.12/site-packages/transformers/tokenization_utils_base.py", line 1825, in get_chat_template [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] raise ValueError( [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] ValueError: Cannot use chat template functions because tokenizer.chat_template is not set and no template argument was passed! For information about writing templates and setting the tokenizer.chat_template attribute, please see the documentation at https://huggingface.co/docs/transformers/main/en/chat_templating [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:141] There is no chat template fallback for /mnt/models [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [renderers/hf.py:314] Detected the chat template content format to be 'string'. You can set `--chat-template-content-format` to override this. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] Failed to load AutoTokenizer chat template for /mnt/models [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] Traceback (most recent call last): [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] File "/opt/venv/lib/python3.12/site-packages/vllm/renderers/hf.py", line 120, in resolve_chat_template [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] return tokenizer.get_chat_template(chat_template, tools=tools) [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] File "/opt/venv/lib/python3.12/site-packages/transformers/tokenization_utils_base.py", line 1825, in get_chat_template [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] raise ValueError( [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/hf.py:122] ValueError: Cannot use chat template functions because tokenizer.chat_template is not set and no template argument was passed! For information about writing templates and setting the tokenizer.chat_template attribute, please see the documentation at https://huggingface.co/docs/transformers/main/en/chat_templating [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:05 [renderers/base.py:205] This model does not support chat template. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/openai/api_server.py:594] Starting vLLM server on https://0.0.0.0:8001 [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:37] Available routes are: [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /openapi.json, Methods: GET, HEAD [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /docs, Methods: GET, HEAD [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /redoc, Methods: GET, HEAD [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /tokenize, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /detokenize, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /load, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /version, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /health, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /metrics, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/models, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /ping, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /ping, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /invocations, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/chat/completions, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/responses, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/completions, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/messages, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /inference/v1/generate, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /scale_elastic_ep, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/chat/completions/render, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/launcher.py:46] Route: /v1/completions/render, Methods: POST [e2e-llm-inference-service] (APIServer pid=1) INFO: Started server process [1] [e2e-llm-inference-service] (APIServer pid=1) INFO: Waiting for application startup. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:05 [entrypoints/ssl.py:60] SSLCertRefresher monitors files: ['/var/run/kserve/tls/tls.key', '/var/run/kserve/tls/tls.crt'] [e2e-llm-inference-service] (APIServer pid=1) INFO: Application startup complete. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:08 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:64] File change detected: modified - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:34] Reloading SSL certificate chain [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:64] File change detected: modified - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:34] Reloading SSL certificate chain [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:64] File change detected: deleted - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:34] Reloading SSL certificate chain [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:64] File change detected: deleted - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 16:57:08 [entrypoints/ssl.py:34] Reloading SSL certificate chain [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:57:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:58:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 16:59:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:00:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:01:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:02:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:03:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:04:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:05:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:06:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:07:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:08:56 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:09:06 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:09:16 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:09:26 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:09:36 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-28 17:09:46 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:172 ### Pod llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b8x7fs (phase=Pending) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:228 # -- logs (current): unavailable ((400) [e2e-llm-inference-service] Reason: Bad Request [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '09316594-cefa-41ef-aba3-79c3fca42a24', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Tue, 28 Jul 2026 17:09:51 GMT', 'Content-Length': '247'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"container \"main\" in pod \"llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b8x7fs\" is waiting to start: ContainerCreating","reason":"BadRequest","code":400} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:172 ### Pod llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:225 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:226 {"level":"info","ts":1785257702.3859324,"logger":"setup","caller":"runner/runner.go:196","msg":"GIE build","commit-sha":"unknown","build-ref":""} [e2e-llm-inference-service] {"level":"info","ts":1785257702.3860407,"logger":"setup","caller":"runner/runner.go:217","msg":"Flags processed","flags":{"cert-path":"/var/run/kserve/tls","config-file":"","config-text":"apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type: disagg-headers-handler\n- type: prefill-filter\n- type: decode-filter\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n- type: active-request-scorer\n- type: prefix-cache-scorer\n- type: max-score-picker\n- type: always-disagg-pd-decider\n- parameters:\n deciders:\n prefill: always-disagg-pd-decider\n type: disagg-profile-handler\n- parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n- name: prefill\n plugins:\n - pluginRef: prefill-filter\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef: max-score-picker\n- name: decode\n plugins:\n - pluginRef: decode-filter\n - pluginRef: active-request-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n","disable-endpoint-subset-filter":false,"enable-cert-reload":true,"enable-grpc-stream-metrics":false,"enable-pprof":true,"endpoint-selector":"","endpoint-target-ports":{},"grpc-health-port":9003,"grpc-max-recv-msg-size":"","grpc-max-send-msg-size":"","grpc-port":9002,"ha-enable-leader-election":false,"health-checking":false,"metrics-endpoint-auth":true,"metrics-port":9090,"metrics-staleness-threshold":2000000000,"model-server-metrics-https-insecure-skip-verify":true,"model-server-metrics-path":"/metrics","model-server-metrics-port":0,"model-server-metrics-scheme":"http","pool-group":"inference.networking.k8s.io","pool-name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","pool-namespace":"e2e-test-llm-inference-service-48639af5","refresh-metrics-interval":50000000,"refresh-prometheus-metrics-interval":5000000000,"secure-serving":true,"tracing":true,"v":2,"zap-devel":{},"zap-encoder":{},"zap-log-level":{},"zap-stacktrace-level":{},"zap-time-encoding":{}}} [e2e-llm-inference-service] {"level":"info","ts":1785257702.3861747,"logger":"setup.trace","caller":"tracing/telemetry.go:123","msg":"init OTel trace exporter","type":"console"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.3868818,"caller":"loader/configloader.go:121","msg":"Loaded raw configuration","config":"{Plugins: [{Type: disagg-headers-handler} {Type: prefill-filter} {Type: decode-filter} {Type: queue-scorer} {Type: kv-cache-utilization-scorer} {Type: active-request-scorer} {Type: prefix-cache-scorer} {Type: max-score-picker} {Type: always-disagg-pd-decider} {Type: disagg-profile-handler, Parameters: {\"deciders\":{\"prefill\":\"always-disagg-pd-decider\"}}} {Type: metrics-data-source, Parameters: {\"scheme\":\"https\"}}], SchedulingProfiles: [{Name: prefill, Plugins: [{PluginRef: prefill-filter} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: queue-scorer, Weight: 2.00} {PluginRef: kv-cache-utilization-scorer, Weight: 2.00} {PluginRef: max-score-picker}]} {Name: decode, Plugins: [{PluginRef: decode-filter} {PluginRef: active-request-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.386914,"logger":"setup","caller":"runner/runner.go:622","msg":"Data layer: ENABLED"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.3872108,"logger":"setup","caller":"runner/runner.go:281","msg":"Raw config after phase one","config":{"apiVersion":"llm-d.ai/v1alpha1","dataLayer":null,"kind":"EndpointPickerConfig","plugins":[{"name":"disagg-headers-handler","parameters":null,"type":"disagg-headers-handler"},{"name":"prefill-filter","parameters":null,"type":"prefill-filter"},{"name":"decode-filter","parameters":null,"type":"decode-filter"},{"name":"queue-scorer","parameters":null,"type":"queue-scorer"},{"name":"kv-cache-utilization-scorer","parameters":null,"type":"kv-cache-utilization-scorer"},{"name":"active-request-scorer","parameters":null,"type":"active-request-scorer"},{"name":"prefix-cache-scorer","parameters":null,"type":"prefix-cache-scorer"},{"name":"max-score-picker","parameters":null,"type":"max-score-picker"},{"name":"always-disagg-pd-decider","parameters":null,"type":"always-disagg-pd-decider"},{"name":"disagg-profile-handler","parameters":{"deciders":{"prefill":"always-disagg-pd-decider"}},"type":"disagg-profile-handler"},{"name":"metrics-data-source","parameters":{"scheme":"https"},"type":"metrics-data-source"}],"schedulingProfiles":[{"name":"prefill","plugins":[{"pluginRef":"prefill-filter","weight":null},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"queue-scorer","weight":2},{"pluginRef":"kv-cache-utilization-scorer","weight":2},{"pluginRef":"max-score-picker","weight":null}]},{"name":"decode","plugins":[{"pluginRef":"decode-filter","weight":null},{"pluginRef":"active-request-scorer","weight":2},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"max-score-picker","weight":null}]}]}} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4083378,"caller":"disagg/disagg_profile_handler.go:186","msg":"No deciders.encode configured, E disaggregation disabled"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.408445,"logger":"utilization-detector/utilization-detector","caller":"utilization/detector.go:83","msg":"Creating new UtilizationDetector","queueDepthThreshold":5,"kvCacheUtilThreshold":0.8,"metricsStalenessThreshold":"200ms","headroom":0} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4085052,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"vllm","mapping":"Mapping{all specs enabled}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4085546,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"sglang","mapping":"Mapping{disabled: [lora]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4085934,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"trtllm-serve","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4086783,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton-tensorrt-llm","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4086998,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton","mapping":"Mapping{disabled: [kv, lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4087794,"caller":"loader/configloader.go:154","msg":"Instantiated all plugins and applied system defaults. Effective raw configuration","config":"{Plugins: [{Name: disagg-headers-handler, Type: disagg-headers-handler} {Name: prefill-filter, Type: prefill-filter} {Name: decode-filter, Type: decode-filter} {Name: queue-scorer, Type: queue-scorer} {Name: kv-cache-utilization-scorer, Type: kv-cache-utilization-scorer} {Name: active-request-scorer, Type: active-request-scorer} {Name: prefix-cache-scorer, Type: prefix-cache-scorer} {Name: max-score-picker, Type: max-score-picker} {Name: always-disagg-pd-decider, Type: always-disagg-pd-decider} {Name: disagg-profile-handler, Type: disagg-profile-handler, Parameters: {\"deciders\":{\"prefill\":\"always-disagg-pd-decider\"}}} {Name: metrics-data-source, Type: metrics-data-source, Parameters: {\"scheme\":\"https\"}} {Name: fcfs-ordering-policy, Type: fcfs-ordering-policy} {Name: global-strict-fairness-policy, Type: global-strict-fairness-policy} {Name: static-usage-limit-policy, Type: static-usage-limit-policy} {Name: openai-parser, Type: openai-parser} {Name: anthropic-parser, Type: anthropic-parser} {Name: vllmhttp-parser, Type: vllmhttp-parser} {Name: utilization-detector, Type: utilization-detector} {Name: core-metrics-extractor, Type: core-metrics-extractor}], SchedulingProfiles: [{Name: prefill, Plugins: [{PluginRef: prefill-filter} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: queue-scorer, Weight: 2.00} {PluginRef: kv-cache-utilization-scorer, Weight: 2.00} {PluginRef: max-score-picker}]} {Name: decode, Plugins: [{PluginRef: decode-filter} {PluginRef: active-request-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}], DataLayer: {Sources: [{PluginRef: metrics-data-source, Extractors: [{PluginRef: core-metrics-extractor}]}], Discovery: }, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4088778,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"inflight-load-producer/inflight-load-producer","dataKey":"InFlightLoadDataKey/inflight-load-producer","consumer":"active-request-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4088964,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4089634,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4089859,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"disagg-profile-handler"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.409003,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"disagg-profile-handler"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.409148,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4091609,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: disagg-profile-handler/disagg-profile-handler, Profiles: map[decode:{Filters: [decode-filter/by-label], Scorers: [active-request-scorer/active-request-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker} prefill:{Filters: [prefill-filter/by-label], Scorers: [prefix-cache-scorer/prefix-cache-scorer: 3.000000, queue-scorer/queue-scorer: 2.000000, kv-cache-utilization-scorer/kv-cache-utilization-scorer: 2.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc000731640 DataConfig:{Sources:[{Plugin:0xc000393950 Extractors:[0xc000731840]}]} FlowControlConfig: ParserRegistry:0xc000731f00}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4284556,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.428498,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285202,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285278,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285345,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285407,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285474,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.428553,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285593,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285662,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.428573,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285793,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4285975,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4286065,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: disagg-profile-handler/disagg-profile-handler, Profiles: map[decode:{Filters: [decode-filter/by-label], Scorers: [active-request-scorer/active-request-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker} prefill:{Filters: [prefill-filter/by-label], Scorers: [prefix-cache-scorer/prefix-cache-scorer: 3.000000, queue-scorer/queue-scorer: 2.000000, kv-cache-utilization-scorer/kv-cache-utilization-scorer: 2.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.428667,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4286807,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4287055,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4287422,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4287572,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4288564,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1785257702.428868,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4289172,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4291124,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4294865,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4295895,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4296045,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4298105,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4299586,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4300616,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4301453,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4304118,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1785257702.437307,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4373012,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.4377487,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.43824,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5311315,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.531172,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5311732,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5311892,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5311465,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5312076,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785257702.5313184,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","namespace":"e2e-test-llm-inference-service-48639af5"},"namespace":"e2e-test-llm-inference-service-48639af5","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","reconcileID":"a0d4fa17-5c52-4899-b8f3-dd701b1096b0"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.631256,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1785257702.6313152,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785257711.0901108,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","namespace":"e2e-test-llm-inference-service-48639af5"},"namespace":"e2e-test-llm-inference-service-48639af5","name":"llmisvc-model-pvc-router-manage-e8706282-inference-pool","reconcileID":"108b622a-b0f1-4503-8552-acb0804caeec"} [e2e-llm-inference-service] {"level":"info","ts":1785257828.2673607,"caller":"inflightload/producer.go:308","msg":"Injected dynamic attribute into endpoint","key":"InFlightLoadDataKey/inflight-load-producer","endpoint":"e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m-rank-0"} [e2e-llm-inference-service] {"level":"info","ts":1785257828.2674055,"caller":"controller/pod_reconciler.go:99","msg":"Pod already exists","controller":"pod","controllerGroup":"","controllerKind":"Pod","Pod":{"name":"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m","namespace":"e2e-test-llm-inference-service-48639af5"},"namespace":"e2e-test-llm-inference-service-48639af5","name":"llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m","reconcileID":"609c9606-1419-4ff3-9d64-544a3cf6b606"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: d8f82499-f771-49f2-b3b7-06a240308cf0 [e2e-llm-inference-service] resourceVersion: '59275' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.132.0.41 [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 [e2e-llm-inference-service] uid: 41143c07-a7d8-446d-acbf-e55e1c1c82cd [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 3e4c16c8-35ec-44b3-9808-b61ab0bdf956 [e2e-llm-inference-service] resourceVersion: '60258' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.133.0.54 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] uid: ae180e1a-2ade-4406-a572-9f18fc3ca404 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56c- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: ae180e1a-2ade-4406-a572-9f18fc3ca404 [e2e-llm-inference-service] resourceVersion: '60256' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] pod-template-hash: 6d9fffb56c [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.54/23"],"mac_address":"0a:58:0a:85:00:36","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.54/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.54\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:36\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56c [e2e-llm-inference-service] uid: fb6e5d27-0b0c-40b8-b6c9-9796591d568f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-139-138 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"fb6e5d27-0b0c-40b8-b6c9-9796591d568f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8001,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"llm-d-routing-sidecar"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAMESPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fieldRef: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.54"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] - name: kube-api-access-dtzz4 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: llm-d-routing-sidecar [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/pd-sidecar [e2e-llm-inference-service] - --port=8000 [e2e-llm-inference-service] - --vllm-port=8001 [e2e-llm-inference-service] - --kv-connector=nixlv2 [e2e-llm-inference-service] - --enable-ssrf-protection=true [e2e-llm-inference-service] - --pool-group=inference.networking.x-k8s.io [e2e-llm-inference-service] - --inference-pool=e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --secure-proxy=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] - --enable-tls=decoder [e2e-llm-inference-service] - --enable-tls=prefiller [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: INFERENCE_POOL_NAMESPACE [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] - name: INFERENCE_POOL_NAME [e2e-llm-inference-service] value: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] resources: {} [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-dtzz4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 10 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000970000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to infer\ [e2e-llm-inference-service] \ RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/* 2>/dev/null\n\ [e2e-llm-inference-service] \ grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/* 2>/dev/null\n\ [e2e-llm-inference-service] \n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"$hca_dir\"\ [e2e-llm-inference-service] \ ]; then\n hca_name=$(basename \"$hca_dir\")\n port_state_file=\"\ [e2e-llm-inference-service] $hca_dir/ports/1/state\" # Assume port 1\n type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\ [e2e-llm-inference-service] \n\n echo \"[Infer RoCE] Check if the port state file ${port_state_file}\ [e2e-llm-inference-service] \ exists and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] &&\ [e2e-llm-inference-service] \ grep -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found active\ [e2e-llm-inference-service] \ HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n else\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Skipping inactive or down HCA: $hca_name\"\ [e2e-llm-inference-service] \n fi\n fi\n done\n\n # Check if we found any active HCAs\n\ [e2e-llm-inference-service] \ if [ ${#active_hcas[@]} -gt 0 ]; then\n # Join the array elements with\ [e2e-llm-inference-service] \ a comma\n hca_port_pairs=()\n for hca in \"${active_hcas[@]}\";\ [e2e-llm-inference-service] \ do\n hca_port_pairs+=(\"${hca}:1\")\n done\n\n active_hca_list=$(IFS=,;\ [e2e-llm-inference-service] \ echo \"${active_hcas[*]}\")\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\"\ [e2e-llm-inference-service] )\n echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n\ [e2e-llm-inference-service] \ export UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found. NCCL_IB_HCA\ [e2e-llm-inference-service] \ will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Finding GID_INDEX for each active HCA (SR-IOV compatible)...\"\ [e2e-llm-inference-service] \n\n # For SR-IOV environments, find the most common IPv4 RoCE v2 GID index\ [e2e-llm-inference-service] \ across all HCAs\n declare -A gid_index_count\n declare -A hca_gid_index\n\ [e2e-llm-inference-service] \n for hca_name in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Processing HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for\ [e2e-llm-inference-service] \ this HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"$tpath\"\ [e2e-llm-inference-service] \ 2>/dev/null; then\n idx=$(basename \"$tpath\")\n \ [e2e-llm-inference-service] \ gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n \ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo \"\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Found IPv4 RoCE v2 GID for ${hca_name}:\ [e2e-llm-inference-service] \ index=${idx}, gid=${gid_value}\"\n hca_gid_index[\"${hca_name}\"\ [e2e-llm-inference-service] ]=\"${idx}\"\n gid_index_count[\"${idx}\"]=$((${gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]} + 1))\n break # Use first found IPv4 GID per\ [e2e-llm-inference-service] \ HCA\n fi\n fi\n done\n done\n\n\ [e2e-llm-inference-service] \ # Find the most common GID index (most likely to be consistent across\ [e2e-llm-inference-service] \ nodes)\n best_gid_index=\"\"\n max_count=0\n for idx in \"\ [e2e-llm-inference-service] ${!gid_index_count[@]}\"; do\n count=${gid_index_count[\"${idx}\"]}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n \ [e2e-llm-inference-service] \ if [ $count -gt $max_count ]; then\n max_count=$count\n\ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n #\ [e2e-llm-inference-service] \ Use deterministic fallback if tied - prefer index 3 (SR-IOV standard)\n \ [e2e-llm-inference-service] \ if [ ${#gid_index_count[@]} -gt 1 ]; then\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Multiple GID indices found, selecting most common: ${best_gid_index}\"\n \ [e2e-llm-inference-service] \ # If there's a tie, prefer index 3 as it's most common in SR-IOV setups\n\ [e2e-llm-inference-service] \ if [ -n \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\"\ [e2e-llm-inference-service] \ -eq \"$max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for NCCL,\ [e2e-llm-inference-service] \ NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR: No valid\ [e2e-llm-inference-service] \ IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any HCA.\"\n \ [e2e-llm-inference-service] \ fi\n else\n echo \"[Infer RoCE] No active HCAs found, skipping GID_INDEX\ [e2e-llm-inference-service] \ inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints landed in vLLM\ [e2e-llm-inference-service] \ 0.16.0 (vllm-project/vllm#30011).\n# Older versions still need the blanket\ [e2e-llm-inference-service] \ --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+ ]] &&\ [e2e-llm-inference-service] \ [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort -V | head\ [e2e-llm-inference-service] \ -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout 40\"\ [e2e-llm-inference-service] \nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM 0.22.0+\ [e2e-llm-inference-service] \ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" != *\"\ [e2e-llm-inference-service] --kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"* ]] &&\ [e2e-llm-inference-service] \ [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\"\ [e2e-llm-inference-service] \n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name \"\ [e2e-llm-inference-service] facebook/opt-125m\" \"publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8001 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8001 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-dtzz4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8001 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000970000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c20 [e2e-llm-inference-service] fsGroup: 1000970000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-kserve-dockercfg-7zb6m [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:08Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:08Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] hostIP: 10.0.139.138 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.139.138 [e2e-llm-inference-service] podIP: 10.133.0.54 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.54 [e2e-llm-inference-service] startTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: llm-d-routing-sidecar [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T16:55:08Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-llm-d-router-disagg-sidecar@sha256:4d57f8fe65a63b44f345487af2b831e64ddf7fb01126d4f51cc8cf49b7412d00 [e2e-llm-inference-service] containerID: cri-o://2ffc863da53909247e23a749d1654e991b92f21e7c3af78f73b28641f85c6e32 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] resources: {} [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-dtzz4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000970000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000970000 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T16:55:08Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: docker.io/vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imageID: docker.io/vllm/vllm-openai-cpu@sha256:a8257c201fd2f696c146615987be61579aa2052367b73c27bd3fb1ab8e90bd4b [e2e-llm-inference-service] containerID: cri-o://57ad3a0cc0b3e0336f36edef413f00956632d32be1883b3d2a5d8eb23709e31c [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-dtzz4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000970000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000970000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b8x7fs [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b89cd7957- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: f426972c-46df-4df2-8acc-1879cc25f8cd [e2e-llm-inference-service] resourceVersion: '58693' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] pod-template-hash: 7b89cd7957 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.71/23"],"mac_address":"0a:58:0a:86:00:47","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.71/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b89cd7957 [e2e-llm-inference-service] uid: a27fc0c7-df15-48c4-8f47-f16ce0d4eb28 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-140-86 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"a27fc0c7-df15-48c4-8f47-f16ce0d4eb28"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] - name: kube-api-access-w2zk8 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to infer\ [e2e-llm-inference-service] \ RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/* 2>/dev/null\n\ [e2e-llm-inference-service] \ grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/* 2>/dev/null\n\ [e2e-llm-inference-service] \n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"$hca_dir\"\ [e2e-llm-inference-service] \ ]; then\n hca_name=$(basename \"$hca_dir\")\n port_state_file=\"\ [e2e-llm-inference-service] $hca_dir/ports/1/state\" # Assume port 1\n type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\ [e2e-llm-inference-service] \n\n echo \"[Infer RoCE] Check if the port state file ${port_state_file}\ [e2e-llm-inference-service] \ exists and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] &&\ [e2e-llm-inference-service] \ grep -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found active\ [e2e-llm-inference-service] \ HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n else\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Skipping inactive or down HCA: $hca_name\"\ [e2e-llm-inference-service] \n fi\n fi\n done\n\n # Check if we found any active HCAs\n\ [e2e-llm-inference-service] \ if [ ${#active_hcas[@]} -gt 0 ]; then\n # Join the array elements with\ [e2e-llm-inference-service] \ a comma\n hca_port_pairs=()\n for hca in \"${active_hcas[@]}\";\ [e2e-llm-inference-service] \ do\n hca_port_pairs+=(\"${hca}:1\")\n done\n\n active_hca_list=$(IFS=,;\ [e2e-llm-inference-service] \ echo \"${active_hcas[*]}\")\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\"\ [e2e-llm-inference-service] )\n echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n\ [e2e-llm-inference-service] \ export UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found. NCCL_IB_HCA\ [e2e-llm-inference-service] \ will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Finding GID_INDEX for each active HCA (SR-IOV compatible)...\"\ [e2e-llm-inference-service] \n\n # For SR-IOV environments, find the most common IPv4 RoCE v2 GID index\ [e2e-llm-inference-service] \ across all HCAs\n declare -A gid_index_count\n declare -A hca_gid_index\n\ [e2e-llm-inference-service] \n for hca_name in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Processing HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for\ [e2e-llm-inference-service] \ this HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"$tpath\"\ [e2e-llm-inference-service] \ 2>/dev/null; then\n idx=$(basename \"$tpath\")\n \ [e2e-llm-inference-service] \ gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n \ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo \"\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Found IPv4 RoCE v2 GID for ${hca_name}:\ [e2e-llm-inference-service] \ index=${idx}, gid=${gid_value}\"\n hca_gid_index[\"${hca_name}\"\ [e2e-llm-inference-service] ]=\"${idx}\"\n gid_index_count[\"${idx}\"]=$((${gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]} + 1))\n break # Use first found IPv4 GID per\ [e2e-llm-inference-service] \ HCA\n fi\n fi\n done\n done\n\n\ [e2e-llm-inference-service] \ # Find the most common GID index (most likely to be consistent across\ [e2e-llm-inference-service] \ nodes)\n best_gid_index=\"\"\n max_count=0\n for idx in \"\ [e2e-llm-inference-service] ${!gid_index_count[@]}\"; do\n count=${gid_index_count[\"${idx}\"]}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n \ [e2e-llm-inference-service] \ if [ $count -gt $max_count ]; then\n max_count=$count\n\ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n #\ [e2e-llm-inference-service] \ Use deterministic fallback if tied - prefer index 3 (SR-IOV standard)\n \ [e2e-llm-inference-service] \ if [ ${#gid_index_count[@]} -gt 1 ]; then\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Multiple GID indices found, selecting most common: ${best_gid_index}\"\n \ [e2e-llm-inference-service] \ # If there's a tie, prefer index 3 as it's most common in SR-IOV setups\n\ [e2e-llm-inference-service] \ if [ -n \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\"\ [e2e-llm-inference-service] \ -eq \"$max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for NCCL,\ [e2e-llm-inference-service] \ NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR: No valid\ [e2e-llm-inference-service] \ IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any HCA.\"\n \ [e2e-llm-inference-service] \ fi\n else\n echo \"[Infer RoCE] No active HCAs found, skipping GID_INDEX\ [e2e-llm-inference-service] \ inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints landed in vLLM\ [e2e-llm-inference-service] \ 0.16.0 (vllm-project/vllm#30011).\n# Older versions still need the blanket\ [e2e-llm-inference-service] \ --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+ ]] &&\ [e2e-llm-inference-service] \ [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort -V | head\ [e2e-llm-inference-service] \ -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout 40\"\ [e2e-llm-inference-service] \nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM 0.22.0+\ [e2e-llm-inference-service] \ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" != *\"\ [e2e-llm-inference-service] --kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"* ]] &&\ [e2e-llm-inference-service] \ [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\"\ [e2e-llm-inference-service] \n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name \"\ [e2e-llm-inference-service] facebook/opt-125m\" \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-w2zk8 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: File [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000970000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-140-86.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c20 [e2e-llm-inference-service] fsGroup: 1000970000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Pending [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] hostIP: 10.0.140.86 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.140.86 [e2e-llm-inference-service] startTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] waiting: [e2e-llm-inference-service] reason: ContainerCreating [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imageID: '' [e2e-llm-inference-service] started: false [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-w2zk8 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 [e2e-llm-inference-service] generateName: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-7fc69df7b8- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 41143c07-a7d8-446d-acbf-e55e1c1c82cd [e2e-llm-inference-service] resourceVersion: '59274' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7fc69df7b8 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.41/23"],"mac_address":"0a:58:0a:84:00:29","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.41/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.41\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:29\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-7fc69df7b8 [e2e-llm-inference-service] uid: e6e0d5b1-cb99-44c8-a774-d4b40d54417d [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-142-94 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e6e0d5b1-cb99-44c8-a774-d4b40d54417d"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.41"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-rbpxb [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type:\ [e2e-llm-inference-service] \ disagg-headers-handler\n- type: prefill-filter\n- type: decode-filter\n- type:\ [e2e-llm-inference-service] \ queue-scorer\n- type: kv-cache-utilization-scorer\n- type: active-request-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: max-score-picker\n- type: always-disagg-pd-decider\n\ [e2e-llm-inference-service] - parameters:\n deciders:\n prefill: always-disagg-pd-decider\n type:\ [e2e-llm-inference-service] \ disagg-profile-handler\n- parameters:\n scheme: https\n type: metrics-data-source\n\ [e2e-llm-inference-service] schedulingProfiles:\n- name: prefill\n plugins:\n - pluginRef: prefill-filter\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: queue-scorer\n\ [e2e-llm-inference-service] \ weight: 2\n - pluginRef: kv-cache-utilization-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: max-score-picker\n- name: decode\n plugins:\n - pluginRef:\ [e2e-llm-inference-service] \ decode-filter\n - pluginRef: active-request-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-rbpxb [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000970000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c20 [e2e-llm-inference-service] fsGroup: 1000970000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-epp-sa-dockercfg-grc8w [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] hostIP: 10.0.142.94 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.142.94 [e2e-llm-inference-service] podIP: 10.132.0.41 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.41 [e2e-llm-inference-service] startTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-llm-d-router-endpoint-picker@sha256:7fd1d61cb5505e026eebab0b9653758779c3dee241e37bea92a0bab961a48c50 [e2e-llm-inference-service] containerID: cri-o://8b89f5707ca9009bbc47b230744c2c728cd821023cb049f9467e17d81bb146f7 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-rbpxb [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000970000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000970000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 52f359aa-4324-4c80-b378-ce99e830b341 [e2e-llm-inference-service] resourceVersion: '58686' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-pvc-router-manage-e8706282-epp-sa-dockercfg-grc8w [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-pvc-router-manage-e8706282-epp-sa-dockercfg-grc8w"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-58g82"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-epp-sa-dockercfg-grc8w [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-epp-sa-dockercfg-grc8w [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 975a67b2-5a9a-4d4f-9cdb-51a4abac359c [e2e-llm-inference-service] resourceVersion: '58645' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-pvc-router-manage-e8706282-kserve-dockercfg-7zb6m [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-pvc-router-manage-e8706282-kserve-dockercfg-7zb6m"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-58g82"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-kserve-dockercfg-7zb6m [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-58g82 [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-e8706282-kserve-dockercfg-7zb6m [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: ef37978a-c459-4382-afa1-4bf0179c4b7f [e2e-llm-inference-service] resourceVersion: '58706' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.124.238 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.124.238 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 5fe6d8bd-aea8-40d3-91b8-f95d414a9da1 [e2e-llm-inference-service] resourceVersion: '58672' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.161.107 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.161.107 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 601efb4d-b5b7-4c30-a6de-d924529a7d16 [e2e-llm-inference-service] resourceVersion: '60262' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8001,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"llm-d-routing-sidecar"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAMESPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fieldRef: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: llm-d-routing-sidecar [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/pd-sidecar [e2e-llm-inference-service] - --port=8000 [e2e-llm-inference-service] - --vllm-port=8001 [e2e-llm-inference-service] - --kv-connector=nixlv2 [e2e-llm-inference-service] - --enable-ssrf-protection=true [e2e-llm-inference-service] - --pool-group=inference.networking.x-k8s.io [e2e-llm-inference-service] - --inference-pool=e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --secure-proxy=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] - --enable-tls=decoder [e2e-llm-inference-service] - --enable-tls=prefiller [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: INFERENCE_POOL_NAMESPACE [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] - name: INFERENCE_POOL_NAME [e2e-llm-inference-service] value: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] resources: {} [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 10 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8001 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8001 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8001 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56c" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 3fce63d7-1f5c-4b9e-a165-f47ca0376546 [e2e-llm-inference-service] resourceVersion: '65056' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:05:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:unavailableReplicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n\ [e2e-llm-inference-service] \ ${SHUTDOWN_TIMEOUT_ARGS} \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh\ [e2e-llm-inference-service] \ \\\n --ssl-certfile /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: File [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] unavailableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:05:02Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:05:02Z' [e2e-llm-inference-service] reason: ProgressDeadlineExceeded [e2e-llm-inference-service] message: ReplicaSet "llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b89cd7957" [e2e-llm-inference-service] has timed out progressing. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 6f11c904-891f-46b8-8759-119b16c6aab2 [e2e-llm-inference-service] resourceVersion: '59278' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: disagg-headers-handler\n- type: prefill-filter\n- type: decode-filter\n\ [e2e-llm-inference-service] - type: queue-scorer\n- type: kv-cache-utilization-scorer\n- type: active-request-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: max-score-picker\n- type: always-disagg-pd-decider\n\ [e2e-llm-inference-service] - parameters:\n deciders:\n prefill: always-disagg-pd-decider\n\ [e2e-llm-inference-service] \ type: disagg-profile-handler\n- parameters:\n scheme: https\n type:\ [e2e-llm-inference-service] \ metrics-data-source\nschedulingProfiles:\n- name: prefill\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: prefill-filter\n - pluginRef: prefix-cache-scorer\n \ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ kv-cache-utilization-scorer\n weight: 2\n - pluginRef: max-score-picker\n\ [e2e-llm-inference-service] - name: decode\n plugins:\n - pluginRef: decode-filter\n - pluginRef:\ [e2e-llm-inference-service] \ active-request-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-7fc69df7b8" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56c [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: fb6e5d27-0b0c-40b8-b6c9-9796591d568f [e2e-llm-inference-service] resourceVersion: '60261' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] pod-template-hash: 6d9fffb56c [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] uid: 601efb4d-b5b7-4c30-a6de-d924529a7d16 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"601efb4d-b5b7-4c30-a6de-d924529a7d16"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8001,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"llm-d-routing-sidecar"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"INFERENCE_POOL_NAMESPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fieldRef: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] pod-template-hash: 6d9fffb56c [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] pod-template-hash: 6d9fffb56c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: llm-d-routing-sidecar [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-disagg-sidecar:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/pd-sidecar [e2e-llm-inference-service] - --port=8000 [e2e-llm-inference-service] - --vllm-port=8001 [e2e-llm-inference-service] - --kv-connector=nixlv2 [e2e-llm-inference-service] - --enable-ssrf-protection=true [e2e-llm-inference-service] - --pool-group=inference.networking.x-k8s.io [e2e-llm-inference-service] - --inference-pool=e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --secure-proxy=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] - --enable-tls=decoder [e2e-llm-inference-service] - --enable-tls=prefiller [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: INFERENCE_POOL_NAMESPACE [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] - name: INFERENCE_POOL_NAME [e2e-llm-inference-service] value: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] resources: {} [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 10 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8001 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8001 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8001 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-7b89cd7957 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: a27fc0c7-df15-48c4-8f47-f16ce0d4eb28 [e2e-llm-inference-service] resourceVersion: '58689' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] pod-template-hash: 7b89cd7957 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] uid: 3fce63d7-1f5c-4b9e-a165-f47ca0376546 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3fce63d7-1f5c-4b9e-a165-f47ca0376546"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] pod-template-hash: 7b89cd7957 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-prefill [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: prefill [e2e-llm-inference-service] pod-template-hash: 7b89cd7957 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n\ [e2e-llm-inference-service] \ ${SHUTDOWN_TIMEOUT_ARGS} \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh\ [e2e-llm-inference-service] \ \\\n --ssl-certfile /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 180 [e2e-llm-inference-service] timeoutSeconds: 30 [e2e-llm-inference-service] periodSeconds: 30 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 8 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 5 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: File [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-7fc69df7b8 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: e6e0d5b1-cb99-44c8-a774-d4b40d54417d [e2e-llm-inference-service] resourceVersion: '59277' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7fc69df7b8 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler [e2e-llm-inference-service] uid: 6f11c904-891f-46b8-8759-119b16c6aab2 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"6f11c904-891f-46b8-8759-119b16c6aab2"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7fc69df7b8 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7fc69df7b8 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: disagg-headers-handler\n- type: prefill-filter\n- type: decode-filter\n\ [e2e-llm-inference-service] - type: queue-scorer\n- type: kv-cache-utilization-scorer\n- type: active-request-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: max-score-picker\n- type: always-disagg-pd-decider\n\ [e2e-llm-inference-service] - parameters:\n deciders:\n prefill: always-disagg-pd-decider\n\ [e2e-llm-inference-service] \ type: disagg-profile-handler\n- parameters:\n scheme: https\n type:\ [e2e-llm-inference-service] \ metrics-data-source\nschedulingProfiles:\n- name: prefill\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: prefill-filter\n - pluginRef: prefix-cache-scorer\n \ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ kv-cache-utilization-scorer\n weight: 2\n - pluginRef: max-score-picker\n\ [e2e-llm-inference-service] - name: decode\n plugins:\n - pluginRef: decode-filter\n - pluginRef:\ [e2e-llm-inference-service] \ active-request-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 9df25133-7242-4549-af2c-ca414f70e409 [e2e-llm-inference-service] resourceVersion: '58700' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: b4c72c9e-5c54-4aa7-ac64-2755df8ccf1e [e2e-llm-inference-service] resourceVersion: '58651' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: ddf3d8c7-59da-432a-abd6-9a8a5d470a16 [e2e-llm-inference-service] resourceVersion: '58698' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 9c724ebd-2197-4af3-92b8-482a28a381f9 [e2e-llm-inference-service] resourceVersion: '58649' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service-n54lw [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-e8706282-epp-service- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 1f12c5b4-995d-48c1-bcce-3c1afba112cf [e2e-llm-inference-service] resourceVersion: '59276' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] uid: ef37978a-c459-4382-afa1-4bf0179c4b7f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:33Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"ef37978a-c459-4382-afa1-4bf0179c4b7f"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.41 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 [e2e-llm-inference-service] uid: 41143c07-a7d8-446d-acbf-e55e1c1c82cd [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-sbvsdx [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: b0d52c0d-cca0-466e-bb1c-e3af777220a7 [e2e-llm-inference-service] resourceVersion: '60257' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] uid: 5fe6d8bd-aea8-40d3-91b8-f95d414a9da1 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:57:08Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5fe6d8bd-aea8-40d3-91b8-f95d414a9da1"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.133.0.54 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] uid: ae180e1a-2ade-4406-a572-9f18fc3ca404 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 9df25133-7242-4549-af2c-ca414f70e409 [e2e-llm-inference-service] resourceVersion: '58700' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-inference-service-48639af5:llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: b4c72c9e-5c54-4aa7-ac64-2755df8ccf1e [e2e-llm-inference-service] resourceVersion: '58651' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-inference-service-48639af5:llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: ddf3d8c7-59da-432a-abd6-9a8a5d470a16 [e2e-llm-inference-service] resourceVersion: '58698' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] uid: 9c724ebd-2197-4af3-92b8-482a28a381f9 [e2e-llm-inference-service] resourceVersion: '58649' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59021' [e2e-llm-inference-service] uid: 93062a35-000a-4722-b614-5d0a2aa33327 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59021' [e2e-llm-inference-service] uid: 93062a35-000a-4722-b614-5d0a2aa33327 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-48639af5/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-inference-service-48639af5/llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59002' [e2e-llm-inference-service] uid: c088c554-ef00-49aa-9ecc-445907f0c400 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:03Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T16:55:05Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58903' [e2e-llm-inference-service] uid: eadb283c-9cf8-48ee-bd6b-669d4d4db790 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:04Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T16:55:05Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58776' [e2e-llm-inference-service] uid: dc75724e-45f6-4b15-9701-2e5d74f5b6eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59009' [e2e-llm-inference-service] uid: 000fc421-a6fb-4ae1-abc0-2e11301d7f38 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-inference--ip-8bbb5aea.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58799' [e2e-llm-inference-service] uid: e690a0ba-efad-49b1-93ec-225a016613eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58776' [e2e-llm-inference-service] uid: dc75724e-45f6-4b15-9701-2e5d74f5b6eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59009' [e2e-llm-inference-service] uid: 000fc421-a6fb-4ae1-abc0-2e11301d7f38 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-inference--ip-8bbb5aea.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58799' [e2e-llm-inference-service] uid: e690a0ba-efad-49b1-93ec-225a016613eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58776' [e2e-llm-inference-service] uid: dc75724e-45f6-4b15-9701-2e5d74f5b6eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-epp-service.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:11Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '59009' [e2e-llm-inference-service] uid: 000fc421-a6fb-4ae1-abc0-2e11301d7f38 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-inference--ip-8bbb5aea.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:02Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58799' [e2e-llm-inference-service] uid: e690a0ba-efad-49b1-93ec-225a016613eb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc.e2e-test-llm-inference-service-48639af5.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e41cb095-4571-406d-8488-55916968f48c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T16:55:01Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] uid: e41cb095-4571-406d-8488-55916968f48c [e2e-llm-inference-service] resourceVersion: '58715' [e2e-llm-inference-service] uid: dde5bf4b-056d-4349-bccf-bcca56e7254f [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-e8706282-kserve-6d9fffb56cr8h5m [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:09:53Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: decode [e2e-llm-inference-service] pod-template-hash: 6d9fffb56c [e2e-llm-inference-service] timestamp: '2026-07-28T17:09:28Z' [e2e-llm-inference-service] window: 19.882s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 99344480n [e2e-llm-inference-service] memory: 2400592Ki [e2e-llm-inference-service] - name: llm-d-routing-sidecar [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 15906297n [e2e-llm-inference-service] memory: 23128Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schedcq46 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-48639af5 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:09:53Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-e8706282 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7fc69df7b8 [e2e-llm-inference-service] timestamp: '2026-07-28T17:09:40Z' [e2e-llm-inference-service] window: 18.47s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 59233946n [e2e-llm-inference-service] memory: 32892Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-28T17:09:53.635495] end - ❌ 902.650s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T16:57:20Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:03Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T16:55:40Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T16:55:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-inference-service-48639af5 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_autoscaling_hpa_lws[router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-hpa] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-hpa'], prompt='KSer... {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.autoscaling_hpa [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator-lws", [e2e-llm-inference-service] "prometheus-scrape", [e2e-llm-inference-service] "scaling-hpa", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="autoscale-hpa-lws", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_autoscaling_hpa_lws(test_case: TestCase): [e2e-llm-inference-service] """HPA + LWS: HPA exists with WVA annotations; pods scale under load.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = _new_kserve_client() [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] ns = test_case.namespace [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _create_and_wait(kserve_client, test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:670: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-hpa'], prompt='KSer... {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_and_wait(kserve_client, test_case): [e2e-llm-inference-service] """Create LLMISVC and wait for it to be ready.""" [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:482: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...ale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:05:35.283367', start_time = 1785258335.2836626 [e2e-llm-inference-service] duration = 900.3671391010284, timestamp_end = '2026-07-28T17:20:35.650804' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....autoscale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f1be0131d00> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-autoscale-hpa-lw-1aa98714 in namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-autoscale-hpa-lw-1aa98714 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-autoscale-hpa-lw-1aa98714 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-fe7a55cc in namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-fe7a55cc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-fe7a55cc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-b29acdba in namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-b29acdba [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig prometheus-scrape-autoscale-hpa-b29acdba [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-lws-b344a3ff in namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-lws-b344a3ff [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig scaling-hpa-autoscale-hpa-lws-b344a3ff [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_autoscaling_hpa_lws] [2026-07-28T17:05:34.990178] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-hpa'], prompt='KServe is a', service_name='autoscale-hpa-lws', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-lw-1aa98714'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-fe7a55cc'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:05:35.003181] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-lw-1aa98714'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-fe7a55cc'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:05:35.283226] end - ✅ in 0.280s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:05:35.283367] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-lw-1aa98714'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-fe7a55cc'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/autoscale-hpa-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'Inference Pool e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/autoscale-hpa-lws-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/autoscale-hpa-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/autoscale-hpa-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'HPA conditions not yet available', 'reason': 'HPAProgressing', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 9: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 10: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 11: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:20:35.650804] end - ❌ 900.367s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T17:20:35.650890] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-hpa-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-hpa-lw-1aa98714'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-fe7a55cc'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-hpa-b29acdba'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-hpa-lws-b344a3ff'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 3 pod(s) for autoscale-hpa-lws still terminating: ['autoscale-hpa-lws-kserve-mn-0', 'autoscale-hpa-lws-kserve-mn-0-1', 'autoscale-hpa-lws-kserve-router-scheduler-5b6578b469-fklfd'] [e2e-llm-inference-service] assert not ['autoscale-hpa-lws-kserve-mn-0', 'autoscale-hpa-lws-kserve-mn-0-1', 'autoscale-hpa-lws-kserve-router-scheduler-5b6578b469-fklfd'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 1 pod(s) for autoscale-hpa-lws still terminating: ['autoscale-hpa-lws-kserve-router-scheduler-5b6578b469-fklfd'] [e2e-llm-inference-service] assert not ['autoscale-hpa-lws-kserve-router-scheduler-5b6578b469-fklfd'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-28T17:21:26.066667] end - ✅ in 50.415s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_autoscaling_hpa_lws] [2026-07-28T17:21:26.066800] end - ❌ 951.076s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:05:56Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:06:38Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:06:12Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-hpa-lws-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-autoscaling-hpa-lws-d4cbcfd2 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_inference_service[router-managed-workload-simulated-dp-ep-cpu-model-pvc] _ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-simulated-dp-ep-cpu', 'model-pvc'], prompt='KServe is a', service_name... {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway="router-gateway-1", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway="router-gateway-1", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] routes=[ [e2e-llm-inference-service] make_router_main_route( [e2e-llm-inference-service] "router-route-1", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] "router-with-refs-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] make_router_health_route( [e2e-llm-inference-service] "router-route-2", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-1", [e2e-llm-inference-service] "router-with-refs-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway="router-gateway-2", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_router_resources( [e2e-llm-inference-service] gateways=[ [e2e-llm-inference-service] make_router_gateway( [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] routes=[ [e2e-llm-inference-service] make_router_main_route( [e2e-llm-inference-service] "router-route-3", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] "router-with-refs-pd-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] make_router_health_route( [e2e-llm-inference-service] "router-route-4", [e2e-llm-inference-service] tc.namespace, [e2e-llm-inference-service] "router-gateway-2", [e2e-llm-inference-service] "router-with-refs-pd-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda tc: create_scheduler_configmap(namespace=tc.namespace) [e2e-llm-inference-service] ], [e2e-llm-inference-service] after_test=[ [e2e-llm-inference-service] lambda tc: delete_scheduler_configmap(namespace=tc.namespace) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Standalone tokenizer — clean path: token-producer in inline config [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-tokenizer-kvcache", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="tokenizer-clean-path-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Standalone tokenizer — migration path: legacy precise-prefix-cache-scorer [e2e-llm-inference-service] # triggers auto-provisioned tokenizer without explicit tokenizer:{} field [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="tokenizer-migration-path-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name="publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "publishers/{namespace}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] "publishers/{namespace}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: "publishers/{namespace}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[lambda tc: ensure_pvc_with_model(namespace=tc.namespace)], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:866: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...pu-ll-699c687c'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:10:25.553681', start_time = 1785258625.553959 [e2e-llm-inference-service] duration = 900.305216550827, timestamp_end = '2026-07-28T17:25:25.859178' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....p-ep-cpu-ll-699c687c'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f2fed1216c0> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-inference-service-7c3633d1 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1959 Created PVC e2e-pvc-model-storage in namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:2099 Created model download Job e2e-pvc-model-download in namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:2124 Model download Job e2e-pvc-model-download completed successfully [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-llmisvc-model-pv-2f105a64 in namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-llmisvc-model-pv-2f105a64 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-llmisvc-model-pv-2f105a64 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-699c687c in namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-699c687c [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-699c687c [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-cfc8d654 in namespace e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-cfc8d654 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig model-pvc-llmisvc-model-pvc-rou-cfc8d654 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-28T17:10:25.192051] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-simulated-dp-ep-cpu', 'model-pvc'], prompt='KServe is a', service_name='llmisvc-model-pvc-router-manage-2577e794', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-inference-service-7c3633d1', before_test=[ at 0x7f2fedd96160>], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-2577e794', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-7c3633d1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-2f105a64'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-699c687c'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:10:25.204631] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-2577e794', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-7c3633d1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-2f105a64'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-699c687c'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:10:25.553540] end - ✅ in 0.349s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:10:25.553681] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-pvc-router-manage-2577e794', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-inference-service-7c3633d1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-pv-2f105a64'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-699c687c'}, [e2e-llm-inference-service] {'name': 'model-pvc-llmisvc-model-pvc-rou-cfc8d654'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'Inference Pool e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:25:25.859178] end - ❌ 900.305s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:903 [router-managed-workload-simulated-dp-ep-cpu-model-pvc] ❌ ERROR: Failed to call llm inference service llmisvc-model-pvc-router-manage-2577e794: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:243 # Diagnostics for 'llmisvc-model-pvc-router-manage-2577e794' in 'e2e-test-llm-inference-service-7c3633d1' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:244 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:247 # LLMInferenceService llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:256 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:25Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:25Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:32Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler-latency-predictor: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tokenizer: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] resourceVersion: '70306' [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-llmisvc-model-pv-2f105a64 [e2e-llm-inference-service] - name: workload-simulated-dp-ep-cpu-ll-699c687c [e2e-llm-inference-service] - name: model-pvc-llmisvc-model-pvc-rou-cfc8d654 [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler-latency-predictor: kserve-config-llm-scheduler-latency-predictor [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tokenizer: kserve-config-llm-tokenizer [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: WorkerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a3278b28eeb9346c89b25d34b00332e3-258552735.us-east-1.elb.amazonaws.com/e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:68 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:69 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-7c3633d1 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-7c3633d1/e2e-pvc-model-download-jfpfg to ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:59 e2e-test-llm-inference-service-7c3633d1 attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-5bcb6d83-5678-415a-ab67-e708a4bf3034" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:07 e2e-test-llm-inference-service-7c3633d1 multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:07 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-142-94.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:07 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-142-94.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:07 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-142-94.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:54 e2e-test-llm-inference-service-7c3633d1 job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-jfpfg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:18 e2e-test-llm-inference-service-7c3633d1 job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:54 e2e-test-llm-inference-service-7c3633d1 persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:54 e2e-test-llm-inference-service-7c3633d1 persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:54 e2e-test-llm-inference-service-7c3633d1 ebs.csi.aws.com_aws-ebs-csi-driver-controller-57f6947994-rxq7b_37543d0b-c2ce-44f2-9725-79e46f1d830a Normal Provisioning External provisioner is provisioning volume for claim "e2e-test-llm-inference-service-7c3633d1/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:09:56 e2e-test-llm-inference-service-7c3633d1 ebs.csi.aws.com_aws-ebs-csi-driver-controller-57f6947994-rxq7b_37543d0b-c2ce-44f2-9725-79e46f1d830a Normal ProvisioningSucceeded Successfully provisioned volume pvc-5bcb6d83-5678-415a-ab67-e708a4bf3034 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-7c3633d1 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-140-86.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:52 e2e-test-llm-inference-service-7c3633d1 attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-5bcb6d83-5678-415a-ab67-e708a4bf3034" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:53 e2e-test-llm-inference-service-7c3633d1 multus Normal AddedInterface Add eth0 [10.134.0.73/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:53 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-140-86.ec2.internal Normal Pulling Pulling image "vllm/vllm-openai-cpu:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:11:23 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-140-86.ec2.internal Normal Pulled Successfully pulled image "vllm/vllm-openai-cpu:v0.19.0" in 30.544s (30.544s including waiting). Image size: 3531178100 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:11:24 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-140-86.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:11:24 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-140-86.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-7c3633d1 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:52 e2e-test-llm-inference-service-7c3633d1 attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-5bcb6d83-5678-415a-ab67-e708a4bf3034" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 leaderworkerset Normal CreatingRevision Creating revision with key 7bc46f896d for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:35 e2e-test-llm-inference-service-7c3633d1 OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.Secret e2e-test-llm-inference-service-7c3633d1/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.RoleBinding e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.ServiceAccount e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.Service e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.ServiceAccount e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /e2e-test-llm-inference-service-5571cf2b7477d60b74b195d7a6a8f03e [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.Role e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created Created v1.RoleBinding e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:11:30 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:21:04 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Updated Updated v1.Secret e2e-test-llm-inference-service-7c3633d1/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:11:30 e2e-test-llm-inference-service-7c3633d1 LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:50 e2e-test-llm-inference-service-7c3633d1 deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-66dcb4d65d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 None e2e-test-llm-inference-service-7c3633d1 Normal Scheduled Successfully assigned e2e-test-llm-inference-service-7c3633d1/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs to ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:51 e2e-test-llm-inference-service-7c3633d1 multus Normal AddedInterface Add eth0 [10.133.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:51 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-139-138.ec2.internal Normal Pulled Container image "quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:51 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-139-138.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:80 2026-07-28 17:10:51 e2e-test-llm-inference-service-7c3633d1 kubelet/ip-10-0-139-138.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:172 ### Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 (phase=Pending) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:228 # -- logs (current): unavailable ((400) [e2e-llm-inference-service] Reason: Bad Request [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '3032ec2a-bd1e-4733-926a-afd29c2529a5', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Tue, 28 Jul 2026 17:25:26 GMT', 'Content-Length': '236'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"container \"main\" in pod \"llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0\" is waiting to start: ContainerCreating","reason":"BadRequest","code":400} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:172 ### Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:225 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:226 INFO 07-28 17:11:36 [importing.py:44] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors. [e2e-llm-inference-service] INFO 07-28 17:11:36 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] █ █ █▄ ▄█ [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.19.0 [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] █▄█▀ █ █ █ █ model /mnt/models [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:299] [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:11:42 [utils.py:233] non-default args: {'model_tag': '/mnt/models', 'ssl_keyfile': '/var/run/kserve/tls/tls.key', 'ssl_certfile': '/var/run/kserve/tls/tls.crt', 'enable_ssl_refresh': True, 'model': '/mnt/models', 'served_model_name': ['facebook/opt-125m']} [e2e-llm-inference-service] (APIServer pid=1) WARNING 07-28 17:11:42 [arg_utils.py:1390] The global random seed is set to 0. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:12:00 [model.py:549] Resolved architecture: OPTForCausalLM [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:12:00 [model.py:1678] Using max model len 2048 [e2e-llm-inference-service] (APIServer pid=1) INFO 07-28 17:12:00 [vllm.py:790] Asynchronous scheduling is enabled. [e2e-llm-inference-service] INFO 07-28 17:12:12 [importing.py:44] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors. [e2e-llm-inference-service] INFO 07-28 17:12:12 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available. [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:17 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/mnt/models', speculative_config=None, tokenizer='/mnt/models', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.float16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=facebook/opt-125m, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'dce': True, 'nan_asserts': False, 'epilogue_fusion': True, 'cpp.dynamic_threads': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:236] auto thread-binding list (id, physical core): [(4, 0), (5, 1), (6, 2), (7, 3)] [e2e-llm-inference-service] [W728 17:12:18.299697729 utils.cpp:76] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_threads_env) [e2e-llm-inference-service] [W728 17:12:18.299724715 utils.cpp:103] Warning: NUMA binding: Using MEMBIND policy for memory allocation on the NUMA nodes (0). Memory allocations will be strictly bound to these NUMA nodes. (function init_cpu_threads_env) [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] OMP threads binding of Process 41: [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] OMP tid: 41, core 4 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] OMP tid: 58, core 5 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] OMP tid: 59, core 6 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] OMP tid: 60, core 7 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_worker.py:109] [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.134.0.73:51661 backend=gloo [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-28 17:12:18 [cpu_model_runner.py:71] Starting to load model /mnt/models... [e2e-llm-inference-service] (EngineCore pid=41) Loading pt checkpoint shards: 0% Completed | 0/1 [00:00}, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.2905214,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1785258651.2905958,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.290655,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"prefix-cache-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.2907085,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"approx-prefix-cache-producer"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.2908573,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.2908716,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, kv-cache-utilization-scorer/kv-cache-utilization-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000, no-hit-lru-scorer/no-hit-lru-scorer: 2.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc000908780 DataConfig:{Sources:[{Plugin:0xc000526ab0 Extractors:[0xc000908980]}]} FlowControlConfig: ParserRegistry:0xc000908ec0}"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3142843,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314323,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143373,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143427,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143473,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143516,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143554,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.31436,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314364,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143687,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314373,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314377,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3143973,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314405,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, kv-cache-utilization-scorer/kv-cache-utilization-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000, no-hit-lru-scorer/no-hit-lru-scorer: 2.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3144367,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314445,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3144588,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314472,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":0} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3144834,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3145726,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3145857,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.314612,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.31482,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3151908,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1785258651.315299,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3155832,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.315577,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3156393,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.315763,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.316172,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3163033,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1785258651.325993,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3260183,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3260496,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.3261485,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"go/pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.417377,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.4174173,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.417443,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"llm-d.ai","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785258651.4174237,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"llm-d.ai","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785258651.417367,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.4174962,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785258651.4176276,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-pvc-router-manage-2577e794-inference-pool","namespace":"e2e-test-llm-inference-service-7c3633d1"},"namespace":"e2e-test-llm-inference-service-7c3633d1","name":"llmisvc-model-pvc-router-manage-2577e794-inference-pool","reconcileID":"7a84caeb-2938-4a10-87cd-087e35a7c32c"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.517727,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1785258651.5177777,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1785258690.5025077,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-pvc-router-manage-2577e794-inference-pool","namespace":"e2e-test-llm-inference-service-7c3633d1"},"namespace":"e2e-test-llm-inference-service-7c3633d1","name":"llmisvc-model-pvc-router-manage-2577e794-inference-pool","reconcileID":"f5162966-139a-408c-a812-9d13e788010a"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 3ab2ef59-b8ed-4a26-a163-184ba5940e3b [e2e-llm-inference-service] resourceVersion: '70189' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.133.0.56 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs [e2e-llm-inference-service] uid: 80eff4d0-5e98-4493-a240-4e1ca0657381 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: a5bc3374-c524-4b51-9f3c-3cde84fa58d1 [e2e-llm-inference-service] resourceVersion: '69328' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: c186ec48-fde5-4a04-8fa8-146b9bdd52ba [e2e-llm-inference-service] resourceVersion: '69357' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '0' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-5cf8fb667d [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: ae449cef3b581b6dbab6ee0440d253929da5cb58 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.44/23"],"mac_address":"0a:58:0a:84:00:2c","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.44/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] openshift.io/scc: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: serviceaccount [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] uid: bc00a932-02a5-44c2-89d0-ff91141f4ad9 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-142-94 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/size: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:apps.kubernetes.io/pod-index: {} [e2e-llm-inference-service] f:controller-revision-hash: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/worker-index: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:statefulset.kubernetes.io/pod-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"bc00a932-02a5-44c2-89d0-ff91141f4ad9"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:hostname: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:subdomain: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] - name: kube-api-access-xpvw9 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0.llmisvc-model-pvc-router-manage-2577e794-kserve-mn.e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-xpvw9 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000990000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c30 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt [e2e-llm-inference-service] hostname: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] subdomain: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Pending [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] hostIP: 10.0.142.94 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.142.94 [e2e-llm-inference-service] startTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] waiting: [e2e-llm-inference-service] reason: ContainerCreating [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imageID: '' [e2e-llm-inference-service] started: false [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-xpvw9 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 656d4f86-149c-4f23-bbdc-d777b2b06876 [e2e-llm-inference-service] resourceVersion: '70209' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '1' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-74696b7c6f [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: ae449cef3b581b6dbab6ee0440d253929da5cb58 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '1' [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.73/23"],"mac_address":"0a:58:0a:86:00:49","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.73/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.134.0.73\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:86:00:49\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/leader-name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] openshift.io/scc: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: serviceaccount [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] uid: ee47808e-6d3b-415c-94ea-f37953197b4c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-140-86 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/leader-name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/size: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:apps.kubernetes.io/pod-index: {} [e2e-llm-inference-service] f:controller-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-index: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-key: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:statefulset.kubernetes.io/pod-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"ee47808e-6d3b-415c-94ea-f37953197b4c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:hostname: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:subdomain: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:53Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:24Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.134.0.73"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] - name: kube-api-access-wc9md [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0.llmisvc-model-pvc-router-manage-2577e794-kserve-mn.e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-wc9md [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000990000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] nodeName: ip-10-0-140-86.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c30 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt [e2e-llm-inference-service] hostname: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] subdomain: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:24Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:24Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:24Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] hostIP: 10.0.140.86 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.140.86 [e2e-llm-inference-service] podIP: 10.134.0.73 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.134.0.73 [e2e-llm-inference-service] startTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T17:11:24Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: docker.io/vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imageID: docker.io/vllm/vllm-openai-cpu@sha256:a8257c201fd2f696c146615987be61579aa2052367b73c27bd3fb1ab8e90bd4b [e2e-llm-inference-service] containerID: cri-o://609979a1a31e03e377d7fe4282960705ad21d2e7a67e2a8a694dc31bde0650f8 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-wc9md [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000990000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs [e2e-llm-inference-service] generateName: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-66dcb4d65d- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 80eff4d0-5e98-4493-a240-4e1ca0657381 [e2e-llm-inference-service] resourceVersion: '70188' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 66dcb4d65d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.56/23"],"mac_address":"0a:58:0a:85:00:38","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.56/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.56\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:38\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-66dcb4d65d [e2e-llm-inference-service] uid: 38dc2563-1fbe-4e2d-9364-f33e44cfb005 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-139-138 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"38dc2563-1fbe-4e2d-9364-f33e44cfb005"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.56"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-gnxs2 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type:\ [e2e-llm-inference-service] \ single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n -\ [e2e-llm-inference-service] \ pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: no-hit-lru-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-gnxs2 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000990000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c31,c30 [e2e-llm-inference-service] fsGroup: 1000990000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e794-epp-sa-dockercfg-hz27z [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] hostIP: 10.0.139.138 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.139.138 [e2e-llm-inference-service] podIP: 10.133.0.56 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.56 [e2e-llm-inference-service] startTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-llm-d-router-endpoint-picker@sha256:7fd1d61cb5505e026eebab0b9653758779c3dee241e37bea92a0bab961a48c50 [e2e-llm-inference-service] containerID: cri-o://f9c115c23d9f5b90d10f86628c429351aebed1ed0ae1a1091089f408527a9eb0 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-gnxs2 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000990000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000990000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: d6c4db5c-5e6a-4d6a-94ce-88b4c74d7bda [e2e-llm-inference-service] resourceVersion: '69334' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-pvc-router-manage-2577e794-epp-sa-dockercfg-hz27z [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-pvc-router-manage-2577e794-epp-sa-dockercfg-hz27z"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-mf7pb"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e794-epp-sa-dockercfg-hz27z [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e794-epp-sa-dockercfg-hz27z [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 3c20b038-0bbe-4363-97ca-5d1b0d63e9aa [e2e-llm-inference-service] resourceVersion: '69316' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-mf7pb"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-mf7pb [e2e-llm-inference-service] - name: llmisvc-model-pvc-router-manage-2577e7-b09f84b1-dockercfg-nhmnt [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 7c59c228-c0e1-41eb-a349-faaf421ea7c4 [e2e-llm-inference-service] resourceVersion: '69379' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.2.237 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.2.237 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 989e8f35-2f32-44c8-b2f8-46219e6b1fca [e2e-llm-inference-service] resourceVersion: '69324' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.121.153 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.121.153 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-74696b7c6f [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 1a0ebd96-e3a4-4a08-bc46-19a6df454f7b [e2e-llm-inference-service] resourceVersion: '69352' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] controller.kubernetes.io/hash: 74696b7c6f [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: ae449cef3b581b6dbab6ee0440d253929da5cb58 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] uid: ee47808e-6d3b-415c-94ea-f37953197b4c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:data: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:controller.kubernetes.io/hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-index: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-key: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"ee47808e-6d3b-415c-94ea-f37953197b4c"}: {} [e2e-llm-inference-service] f:revision: {} [e2e-llm-inference-service] data: [e2e-llm-inference-service] spec: [e2e-llm-inference-service] template: [e2e-llm-inference-service] $patch: replace [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/leader-name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: ae449cef3b581b6dbab6ee0440d253929da5cb58 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] revision: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ControllerRevision [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-5cf8fb667d [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 62633d07-b58c-4bb7-8df9-30783077c942 [e2e-llm-inference-service] resourceVersion: '69331' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] controller.kubernetes.io/hash: 5cf8fb667d [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/replicas: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] uid: bc00a932-02a5-44c2-89d0-ff91141f4ad9 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:data: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/replicas: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:controller.kubernetes.io/hash: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/worker-index: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"bc00a932-02a5-44c2-89d0-ff91141f4ad9"}: {} [e2e-llm-inference-service] f:revision: {} [e2e-llm-inference-service] data: [e2e-llm-inference-service] spec: [e2e-llm-inference-service] template: [e2e-llm-inference-service] $patch: replace [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] revision: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ControllerRevision [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 19d0a78c-c98e-419f-838c-bdd873db404c [e2e-llm-inference-service] resourceVersion: '70192' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n weight: 3\n - pluginRef: no-hit-lru-scorer\n\ [e2e-llm-inference-service] \ weight: 2\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-66dcb4d65d" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-66dcb4d65d [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 38dc2563-1fbe-4e2d-9364-f33e44cfb005 [e2e-llm-inference-service] resourceVersion: '70191' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 66dcb4d65d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler [e2e-llm-inference-service] uid: 19d0a78c-c98e-419f-838c-bdd873db404c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"19d0a78c-c98e-419f-838c-bdd873db404c"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 66dcb4d65d [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 66dcb4d65d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n weight: 3\n - pluginRef: no-hit-lru-scorer\n\ [e2e-llm-inference-service] \ weight: 2\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: dc380df4-5780-49a3-80fb-2b39a34ee006 [e2e-llm-inference-service] resourceVersion: '69373' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: bc3f0766-6a5a-4a5d-9421-776764ec1668 [e2e-llm-inference-service] resourceVersion: '69312' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-prefill [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: ClusterRole [e2e-llm-inference-service] name: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: de2823d2-e1ca-4f02-ae55-2e2c29afb3c9 [e2e-llm-inference-service] resourceVersion: '69370' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service-jr2hd [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-2577e794-epp-service- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 16166423-ceee-4f2d-b4b3-6c21588e1a76 [e2e-llm-inference-service] resourceVersion: '70190' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] uid: 7c59c228-c0e1-41eb-a349-faaf421ea7c4 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:11:22Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"7c59c228-c0e1-41eb-a349-faaf421ea7c4"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.133.0.56 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs [e2e-llm-inference-service] uid: 80eff4d0-5e98-4493-a240-4e1ca0657381 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-sd7dbq [e2e-llm-inference-service] generateName: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc- [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: 3df17c0c-8f52-4164-bd0d-303b5e0b53ad [e2e-llm-inference-service] resourceVersion: '69327' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] uid: 989e8f35-2f32-44c8-b2f8-46219e6b1fca [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"989e8f35-2f32-44c8-b2f8-46219e6b1fca"}: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: null [e2e-llm-inference-service] ports: null [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: dc380df4-5780-49a3-80fb-2b39a34ee006 [e2e-llm-inference-service] resourceVersion: '69373' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-inference-service-7c3633d1:llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: bc3f0766-6a5a-4a5d-9421-776764ec1668 [e2e-llm-inference-service] resourceVersion: '69312' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-inference-service-7c3633d1:llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-inference-service-7c3633d1:llmisvc-model-pvc-router-manage-2577e794-kserve-mn-prefill [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-prefill [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] name: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] uid: de2823d2-e1ca-4f02-ae55-2e2c29afb3c9 [e2e-llm-inference-service] resourceVersion: '69370' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70304' [e2e-llm-inference-service] uid: 6692b24c-2eec-4dc1-aae9-50c22d6ed35b [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:31Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70304' [e2e-llm-inference-service] uid: 6692b24c-2eec-4dc1-aae9-50c22d6ed35b [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-inference-service-7c3633d1/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-inference-service-7c3633d1/llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70287' [e2e-llm-inference-service] uid: ecb23c3f-dac1-47f9-af77-4f930fce8a1e [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:52Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:10:54Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69566' [e2e-llm-inference-service] uid: 2535fa17-6b66-4fe2-a604-88242de1735a [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:53Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:54Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] kind: LeaderWorkerSet [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:hpaPodSelector: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] manager: lws [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] - apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderWorkerTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:size: {} [e2e-llm-inference-service] f:workerTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-pvc-source"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:persistentVolumeClaim: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:claimName: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:rolloutStrategy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupPolicy: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69368' [e2e-llm-inference-service] uid: de6c7225-593b-4ad6-b237-522e884db858 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] leaderWorkerTemplate: [e2e-llm-inference-service] leaderTemplate: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] restartPolicy: RecreateGroupOnPodRestart [e2e-llm-inference-service] size: 2 [e2e-llm-inference-service] workerTemplate: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-pvc-source [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] serviceAccountName: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] - name: kserve-pvc-source [e2e-llm-inference-service] persistentVolumeClaim: [e2e-llm-inference-service] claimName: e2e-pvc-model-storage [e2e-llm-inference-service] networkConfig: [e2e-llm-inference-service] subdomainPolicy: Shared [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] rolloutStrategy: [e2e-llm-inference-service] rollingUpdateConfiguration: [e2e-llm-inference-service] maxSurge: 0 [e2e-llm-inference-service] maxUnavailable: 1 [e2e-llm-inference-service] partition: 0 [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] startupPolicy: LeaderCreated [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] message: Replicas are progressing [e2e-llm-inference-service] reason: GroupsProgressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Progressing [e2e-llm-inference-service] hpaPodSelector: leaderworkerset.sigs.k8s.io/name=llmisvc-model-pvc-router-manage-2577e794-kserve-mn,leaderworkerset.sigs.k8s.io/worker-index=0 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69416' [e2e-llm-inference-service] uid: 07a08bcb-e47e-440c-b9d7-d65d2a37e877 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70294' [e2e-llm-inference-service] uid: 55987005-0646-40bc-8708-50c169db6123 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-inference--ip-950f33a3.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69420' [e2e-llm-inference-service] uid: 4e9e7a4c-b069-450a-a0d1-ff2ef48761bf [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69416' [e2e-llm-inference-service] uid: 07a08bcb-e47e-440c-b9d7-d65d2a37e877 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70294' [e2e-llm-inference-service] uid: 55987005-0646-40bc-8708-50c169db6123 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-inference--ip-950f33a3.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69420' [e2e-llm-inference-service] uid: 4e9e7a4c-b069-450a-a0d1-ff2ef48761bf [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69416' [e2e-llm-inference-service] uid: 07a08bcb-e47e-440c-b9d7-d65d2a37e877 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-epp-service.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:11:30Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '70294' [e2e-llm-inference-service] uid: 55987005-0646-40bc-8708-50c169db6123 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-inference--ip-950f33a3.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:51Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69420' [e2e-llm-inference-service] uid: 4e9e7a4c-b069-450a-a0d1-ff2ef48761bf [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc.e2e-test-llm-inference-service-7c3633d1.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5366d258-3532-4c28-87fe-747c3023bef6"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:10:50Z' [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] uid: 5366d258-3532-4c28-87fe-747c3023bef6 [e2e-llm-inference-service] resourceVersion: '69391' [e2e-llm-inference-service] uid: ec84461f-2086-4300-947e-cd313093075a [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:25:27Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '1' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-74696b7c6f [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: ae449cef3b581b6dbab6ee0440d253929da5cb58 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7bc46f896d [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '1' [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] timestamp: '2026-07-28T17:25:01Z' [e2e-llm-inference-service] window: 11.102s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 4133669n [e2e-llm-inference-service] memory: 2424088Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schevvccs [e2e-llm-inference-service] namespace: e2e-test-llm-inference-service-7c3633d1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:25:27Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-pvc-router-manage-2577e794 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 66dcb4d65d [e2e-llm-inference-service] timestamp: '2026-07-28T17:25:01Z' [e2e-llm-inference-service] window: 10.879s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 1393418n [e2e-llm-inference-service] memory: 20304Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-28T17:25:27.373403] end - ❌ 902.181s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:11:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:10:52Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-inference-service-7c3633d1 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_autoscaling_keda_lws[router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-keda] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-keda'], prompt='KSe... {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.autoscaling_keda [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator-lws", [e2e-llm-inference-service] "prometheus-scrape", [e2e-llm-inference-service] "scaling-keda", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="autoscale-keda-lws", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_autoscaling_keda_lws(test_case: TestCase): [e2e-llm-inference-service] """KEDA + LWS: ScaledObject exists with WVA annotations; pods scale under load.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = _new_kserve_client() [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] ns = test_case.namespace [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _create_and_wait(kserve_client, test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:728: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-keda'], prompt='KSe... {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_and_wait(kserve_client, test_case): [e2e-llm-inference-service] """Create LLMISVC and wait for it to be ready.""" [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:482: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...e-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:21:27.949626', start_time = 1785259287.9498956 [e2e-llm-inference-service] duration = 900.4395189285278, timestamp_end = '2026-07-28T17:36:28.389417' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....toscale-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f1be0132520> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-autoscaling-keda-lws-e541a132 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-autoscale-keda-l-78828c4a in namespace e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-autoscale-keda-l-78828c4a [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-autoscale-keda-l-78828c4a [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-1696d0b7 in namespace e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-1696d0b7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-lws-aut-1696d0b7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-231d315d in namespace e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-231d315d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig prometheus-scrape-autoscale-ked-231d315d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig scaling-keda-autoscale-keda-lws-1337f511 in namespace e2e-test-llm-autoscaling-keda-lws-e541a132 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig scaling-keda-autoscale-keda-lws-1337f511 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig scaling-keda-autoscale-keda-lws-1337f511 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_autoscaling_keda_lws] [2026-07-28T17:21:27.487126] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator-lws', 'prometheus-scrape', 'scaling-keda'], prompt='KServe is a', service_name='autoscale-keda-lws', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-autoscaling-keda-lws-e541a132', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-lws-e541a132', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-l-78828c4a'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-1696d0b7'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:21:27.499987] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-lws-e541a132', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-l-78828c4a'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-1696d0b7'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:21:27.949484] end - ✅ in 0.449s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:21:27.949626] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-lws-e541a132', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-l-78828c4a'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-1696d0b7'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 12: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-lws-e541a132/autoscale-keda-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'Inference Pool e2e-test-llm-autoscaling-keda-lws-e541a132/autoscale-keda-lws-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-lws-e541a132/autoscale-keda-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-keda-lws-e541a132/autoscale-keda-lws-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'ScaledObject conditions not yet available', 'reason': 'ScaledObjectProgressing', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 13: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 14: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:36:28.389417] end - ❌ 900.440s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T17:36:28.389500] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-keda-lws', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-keda-lws-e541a132', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-keda-l-78828c4a'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-lws-aut-1696d0b7'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-ked-231d315d'}, [e2e-llm-inference-service] {'name': 'scaling-keda-autoscale-keda-lws-1337f511'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 3 pod(s) for autoscale-keda-lws still terminating: ['autoscale-keda-lws-kserve-mn-0', 'autoscale-keda-lws-kserve-mn-0-1', 'autoscale-keda-lws-kserve-router-scheduler-55c6bcf4c9-z264w'] [e2e-llm-inference-service] assert not ['autoscale-keda-lws-kserve-mn-0', 'autoscale-keda-lws-kserve-mn-0-1', 'autoscale-keda-lws-kserve-router-scheduler-55c6bcf4c9-z264w'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 1 pod(s) for autoscale-keda-lws still terminating: ['autoscale-keda-lws-kserve-router-scheduler-55c6bcf4c9-z264w'] [e2e-llm-inference-service] assert not ['autoscale-keda-lws-kserve-router-scheduler-55c6bcf4c9-z264w'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-28T17:36:53.658677] end - ✅ in 25.269s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_autoscaling_keda_lws] [2026-07-28T17:36:53.658816] end - ❌ 926.171s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:22:21Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:23:01Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'severity': 'Info', 'status': 'True', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:22:45Z', 'message': 'failed to ensure HPA is correctly created for ScaledObject: error parsing prometheus metadata: error parsing prometheus metadata: bearer token= is required when bearer auth is enabled', 'reason': 'ScaledObjectCheckFailed', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-autoscaling-keda-lws-e541a132 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_stop_feature[router-managed-workload-single-cpu-model-fb-opt-125m] __ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=... {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="stop-feature-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_stop_feature(test_case: TestCase): [e2e-llm-inference-service] """Test that stopping an LLMInferenceService sets the Ready condition to False with reason Stopped.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Create the service [e2e-llm-inference-service] print(f"Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] # Wait for the service to be ready [e2e-llm-inference-service] print(f"Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] print(f"✅ LLMInferenceService {service_name} is ready") [e2e-llm-inference-service] [e2e-llm-inference-service] # Stop the service by adding the stop annotation [e2e-llm-inference-service] print(f"Stopping LLMInferenceService {service_name}") [e2e-llm-inference-service] stop_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] # Wait for the service to be marked as stopped [e2e-llm-inference-service] print(f"Waiting for LLMInferenceService {service_name} to be stopped") [e2e-llm-inference-service] wait_for_llm_isvc_stopped( [e2e-llm-inference-service] kserve_client, test_case.llm_service, timeout_seconds=120 [e2e-llm-inference-service] ) [e2e-llm-inference-service] print(f"✅ LLMInferenceService {service_name} is stopped") [e2e-llm-inference-service] [e2e-llm-inference-service] # Verify the workload resources are deleted [e2e-llm-inference-service] print(f"Verifying workload resources are deleted for {service_name}") [e2e-llm-inference-service] verify_workload_resources_deleted( [e2e-llm-inference-service] kserve_client, test_case.llm_service, timeout_seconds=120 [e2e-llm-inference-service] ) [e2e-llm-inference-service] print(f"✅ Workload resources deleted for {service_name}") [e2e-llm-inference-service] [e2e-llm-inference-service] # Restart the service by removing the stop annotation [e2e-llm-inference-service] print(f"Restarting LLMInferenceService {service_name}") [e2e-llm-inference-service] restart_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] # Wait for the service to be ready again [e2e-llm-inference-service] print(f"Waiting for LLMInferenceService {service_name} to be ready again") [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py:109: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:30:32.450930', start_time = 1785259832.451243 [e2e-llm-inference-service] duration = 900.4153442382812, timestamp_end = '2026-07-28T17:45:32.866590' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f2fed444f40> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-stop-feature-d9eef2a1 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-stop-feature-tes-eeaa3ea3 in namespace e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-stop-feature-tes-eeaa3ea3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-stop-feature-tes-eeaa3ea3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-single-cpu-stop-featur-e7f09208 in namespace e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-stop-featur-e7f09208 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-stop-featur-e7f09208 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig model-fb-opt-125m-stop-feature-8f213f2f in namespace e2e-test-llm-stop-feature-d9eef2a1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-stop-feature-8f213f2f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-stop-feature-8f213f2f [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_stop_feature] [2026-07-28T17:25:54.076362] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name='stop-feature-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-stop-feature-d9eef2a1', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:25:54.089080] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:25:54.257860] end - ✅ in 0.169s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:25:54.258005] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:26:47Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-stop-feature-d9eef2a1/stop-feature-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'message': 'Inference Pool e2e-test-llm-stop-feature-d9eef2a1/stop-feature-test-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-stop-feature-d9eef2a1/stop-feature-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-stop-feature-d9eef2a1/stop-feature-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-28T17:30:18.952344] end - ✅ in 264.694s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [stop_llmisvc] [2026-07-28T17:30:18.952501] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [stop_llmisvc] [2026-07-28T17:30:19.452489] end - ✅ in 0.500s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_stopped] [2026-07-28T17:30:19.452633] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={'timeout_seconds': 120} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Ready condition status is not False: True [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_stopped] [2026-07-28T17:30:31.752889] end - ✅ in 12.300s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [verify_workload_resources_deleted] [2026-07-28T17:30:31.753175] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={'timeout_seconds': 120} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [verify_workload_resources_deleted] [2026-07-28T17:30:31.820681] end - ✅ in 0.067s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [restart_llmisvc] [2026-07-28T17:30:31.820796] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [restart_llmisvc] [2026-07-28T17:30:32.450766] end - ✅ in 0.629s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:30:32.450930] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'stop-feature-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-stop-feature-d9eef2a1', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-stop-feature-tes-eeaa3ea3'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-stop-featur-e7f09208'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-stop-feature-8f213f2f'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Service is stopped', 'reason': 'Stopped', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Service is stopped', 'reason': 'Stopped', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Service is stopped', 'reason': 'Stopped', 'severity': 'Info', 'status': 'False', 'type': 'PrefillWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Failed to reconcile scheduler: failed to build expected tokenizer deployment: failed to attach model artifacts to tokenizer deployment: secrets "stop-feature-test-epp-sa-dockercfg-w56db" not found', 'reason': 'SchedulerReconcileError', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Failed to reconcile scheduler: failed to build expected tokenizer deployment: failed to attach model artifacts to tokenizer deployment: secrets "stop-feature-test-epp-sa-dockercfg-w56db" not found', 'reason': 'SchedulerReconcileError', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Failed to reconcile scheduler: failed to build expected tokenizer deployment: failed to attach model artifacts to tokenizer deployment: secrets "stop-feature-test-epp-sa-dockercfg-w56db" not found', 'reason': 'SchedulerReconcileError', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Service is stopped', 'reason': 'Stopped', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:30:30Z', 'message': 'Service is stopped', 'reason': 'Stopped', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:45:32.866590] end - ❌ 900.415s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_stop_feature] [2026-07-28T17:45:32.866781] end - ❌ 1178.790s: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:27:31Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:26:47Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:31:15Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:31:50Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-stop-feature-d9eef2a1 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_autoscaling_cleanup_hpa[router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prom... {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.autoscaling_hpa [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator-no-replicas", [e2e-llm-inference-service] "prometheus-scrape", [e2e-llm-inference-service] "scaling-hpa", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="autoscale-cleanup-hpa", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_autoscaling_cleanup_hpa(test_case: TestCase): [e2e-llm-inference-service] """Removing scaling config should delete HPA.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = _new_kserve_client() [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] ns = test_case.namespace [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _create_and_wait(kserve_client, test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:892: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prom... {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_and_wait(kserve_client, test_case): [e2e-llm-inference-service] """Create LLMISVC and wait for it to be ready.""" [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_autoscaling_wva.py:482: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...e-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:36:54.949755', start_time = 1785260214.9500437 [e2e-llm-inference-service] duration = 900.8542320728302, timestamp_end = '2026-07-28T17:51:55.804280' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....toscale-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f1be0133060> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1371: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-autoscale-cleanu-e5a6b97f in namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-autoscale-cleanu-e5a6b97f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-autoscale-cleanu-e5a6b97f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-llmd-simulator-no-repl-5d16e76d in namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-no-repl-5d16e76d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-no-repl-5d16e76d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig prometheus-scrape-autoscale-cle-5a67f5d1 in namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig prometheus-scrape-autoscale-cle-5a67f5d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig prometheus-scrape-autoscale-cle-5a67f5d1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig scaling-hpa-autoscale-cleanup-h-aa1ae037 in namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig scaling-hpa-autoscale-cleanup-h-aa1ae037 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig scaling-hpa-autoscale-cleanup-h-aa1ae037 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_autoscaling_cleanup_hpa] [2026-07-28T17:36:54.599446] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator-no-replicas', 'prometheus-scrape', 'scaling-hpa'], prompt='KServe is a', service_name='autoscale-cleanup-hpa', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-cleanup-hpa', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-cleanu-e5a6b97f'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-5d16e76d'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:36:54.612331] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-cleanup-hpa', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-cleanu-e5a6b97f'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-5d16e76d'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:36:54.949602] end - ✅ in 0.337s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:36:54.949755] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-cleanup-hpa', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-cleanu-e5a6b97f'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-5d16e76d'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 15: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/autoscale-cleanup-hpa-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'Inference Pool e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/autoscale-cleanup-hpa-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/autoscale-cleanup-hpa-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/autoscale-cleanup-hpa-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'HPA conditions not yet available', 'reason': 'HPAProgressing', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 16: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 17: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1391 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:51:55.804280] end - ❌ 900.854s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T17:51:55.804464] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'autoscale-cleanup-hpa', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-autoscale-cleanu-e5a6b97f'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-no-repl-5d16e76d'}, [e2e-llm-inference-service] {'name': 'prometheus-scrape-autoscale-cle-5a67f5d1'}, [e2e-llm-inference-service] {'name': 'scaling-hpa-autoscale-cleanup-h-aa1ae037'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 2 pod(s) for autoscale-cleanup-hpa still terminating: ['autoscale-cleanup-hpa-kserve-6cfdf55fc9-2vxgx', 'autoscale-cleanup-hpa-kserve-router-scheduler-6694c4458b-gt84k'] [e2e-llm-inference-service] assert not ['autoscale-cleanup-hpa-kserve-6cfdf55fc9-2vxgx', 'autoscale-cleanup-hpa-kserve-router-scheduler-6694c4458b-gt84k'] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: 1 pod(s) for autoscale-cleanup-hpa still terminating: ['autoscale-cleanup-hpa-kserve-router-scheduler-6694c4458b-gt84k'] [e2e-llm-inference-service] assert not ['autoscale-cleanup-hpa-kserve-router-scheduler-6694c4458b-gt84k'] [e2e-llm-inference-service] INFO common.gateway_proxy_istio:gateway_proxy_istio.py:421 Snapshot 18: restarts={"kserve-ci-e2e-test/router-gateway-1-openshift-default": "error", "kserve-ci-e2e-test/router-gateway-2-openshift-default": "error", "openshift-ingress/openshift-ai-inference-openshift-default": "1\t{\"running\":{\"startedAt\":\"2026-07-28T16:21:55Z\"}}"} resources={"httproutes": 0, "llminferenceservices": 0, "authpolicies": 0} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-28T17:52:36.110257] end - ✅ in 40.305s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_autoscaling_cleanup_hpa] [2026-07-28T17:52:36.110426] end - ❌ 941.511s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:25Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'severity': 'Info', 'status': 'False', 'type': 'ScalingReady'}, {'lastTransitionTime': '2026-07-28T17:38:17Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:37:48Z', 'message': 'the HPA was unable to compute the replica count: unable to get external metric e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40/wva_desired_replicas/&LabelSelector{MatchLabels:map[string]string{variant_name: autoscale-cleanup-hpa-kserve-hpa,},MatchExpressions:[]LabelSelectorRequirement{},}: unable to fetch metrics from external metrics API: scaledObject name is not specified', 'reason': 'FailedGetExternalMetric', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-autoscaling-cleanup-hpa-a41f0e40 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] _ test_llm_tls_resources[router-managed-workload-single-cpu-model-fb-opt-125m] _ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] name = 'tls-verification-test' [e2e-llm-inference-service] namespace = 'e2e-test-llm-tls-resources-1d9e6e79', version = 'v1alpha1' [e2e-llm-inference-service] [e2e-llm-inference-service] def get_llmisvc( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] name, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] version=constants.KSERVE_V1ALPHA1_VERSION, [e2e-llm-inference-service] ): [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICE, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1204: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-llm-tls-resources-1d9e6e79' [e2e-llm-inference-service] plural = 'llminferenceservices', name = 'tls-verification-test' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1632: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'e2e-test-llm-tls-resources-1d9e6e79' [e2e-llm-inference-service] plural = 'llminferenceservices', name = 'tls-verification-test' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'tls-verification-test', 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', 'plural': 'llminferenceservices', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'name' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method get_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'name' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('name' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['name'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `name` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] if 'name' in local_var_params: [e2e-llm-inference-service] path_params['name'] = local_var_params['name'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1739: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'tls-verification-test', 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', 'plural': 'llminferenceservices', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-llm-tls-resources-1d9e6e79/llminferenceservices/tls-verification-test' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'e2e-test-llm-tls-resources-1d9e6e79'), ('plural', 'llminferenceservices'), ('name', 'tls-verification-test')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-llm-tls-resources-1d9e6e79/llminferenceservices/tls-verification-test' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-llm-tls-resources-1d9e6e79/llminferenceservices/tls-verification-test' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://abf1c34d0e83e43e1b8ecf242edd0ab6-aba9c8a062e7f30b.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/e2e-test-llm-tls-resources-1d9e6e79/llminferenceservices/tls-verification-test' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': '55f53948-bd64-47b0-a3f6-9ba965c020c6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=... {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] service_name="tls-verification-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_tls_resources(test_case: TestCase): [e2e-llm-inference-service] """Verify that TLS-related resources (DestinationRules, cert secrets, service port) [e2e-llm-inference-service] are correctly present or absent based on the enableLLMInferenceServiceTLS flag.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] tls_enabled = _is_tls_enabled() [e2e-llm-inference-service] logger.info(f"enableLLMInferenceServiceTLS = {tls_enabled}") [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_tls.py:133: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...rific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-28T17:52:59.849641', start_time = 1785261179.8499105 [e2e-llm-inference-service] duration = 127.01219582557678, timestamp_end = '2026-07-28T17:55:06.862109' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] all_condition_types = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] ctype = condition.get("type") [e2e-llm-inference-service] all_condition_types.add(ctype) [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(ctype) [e2e-llm-inference-service] [e2e-llm-inference-service] # When TokenizerReady is present, it must also be True [e2e-llm-inference-service] if "TokenizerReady" in all_condition_types: [e2e-llm-inference-service] expected_true_conditions.add("TokenizerReady") [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1376: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7f2fed444360> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1387: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] > out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1339: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] name = 'tls-verification-test' [e2e-llm-inference-service] namespace = 'e2e-test-llm-tls-resources-1d9e6e79', version = 'v1alpha1' [e2e-llm-inference-service] [e2e-llm-inference-service] def get_llmisvc( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] name, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] version=constants.KSERVE_V1ALPHA1_VERSION, [e2e-llm-inference-service] ): [e2e-llm-inference-service] try: [e2e-llm-inference-service] return kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICE, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except client.rest.ApiException as e: [e2e-llm-inference-service] > raise RuntimeError( [e2e-llm-inference-service] f"❌ Exception when calling CustomObjectsApi->" [e2e-llm-inference-service] f"get_namespaced_custom_object for LLMInferenceService: {e}" [e2e-llm-inference-service] ) from e [e2e-llm-inference-service] E RuntimeError: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': '55f53948-bd64-47b0-a3f6-9ba965c020c6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1212: RuntimeError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:82 Created test namespace e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret seaweedfs-s3-creds from kserve-ci-e2e-test to e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:131 Copied secret storage-config from kserve-ci-e2e-test to e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:178 Patched default SA in e2e-test-llm-tls-resources-1d9e6e79 with secret seaweedfs-s3-creds [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:162 ConfigMap odh-kserve-custom-ca-bundle already exists in e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.namespace:namespace.py:159 Copied ConfigMap odh-trusted-ca-bundle from kserve-ci-e2e-test to e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig router-managed-tls-verification-86160844 in namespace e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig router-managed-tls-verification-86160844 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig router-managed-tls-verification-86160844 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 in namespace e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1718 Checking LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 in namespace e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1744 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1754 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_tls_resources] [2026-07-28T17:52:59.387558] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name='tls-verification-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=, response_assertion=.response_assertion at 0x7f2feddce7a0>, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, namespace='e2e-test-llm-tls-resources-1d9e6e79', before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1769 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:test_llm_tls.py:122 enableLLMInferenceServiceTLS = True [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-28T17:52:59.435347] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-28T17:52:59.849495] end - ✅ in 0.414s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-28T17:52:59.849641] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:53:59Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'message': 'Inference Pool e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'message': 'The following HTTPRoutes are not ready: [e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:54:27Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1394 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-28T17:54:27Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:53:59Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-28T17:54:52Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-28T17:54:52Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-28T17:54:27Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-28T17:55:06.862109] end - ❌ 127.012s: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '55f53948-bd64-47b0-a3f6-9ba965c020c6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.test_llm_tls:test_llm_tls.py:141 Failed TLS verification for tls-verification-test: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '55f53948-bd64-47b0-a3f6-9ba965c020c6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:243 # Diagnostics for 'tls-verification-test' in 'e2e-test-llm-tls-resources-1d9e6e79' [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:244 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:247 # LLMInferenceService tls-verification-test [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:258 # failed to dump LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'dfec1896-7ed3-41d4-a93a-f7991ebedd3e', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:68 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:69 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 None e2e-test-llm-tls-resources-1d9e6e79 Normal Scheduled Successfully assigned e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-784875bf86-2bg2p to ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:02 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Pulled Container image "vllm/vllm-openai-cpu:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:02 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:02 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-142-94.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8000/health": dial tcp 10.132.0.50:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-784875bf86-2bg2p [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 None e2e-test-llm-tls-resources-1d9e6e79 Normal Scheduled Successfully assigned e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-router-scheduler-5585754765-zj4ms to ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 multus Normal AddedInterface Add eth0 [10.133.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-139-138.ec2.internal Normal Pulled Container image "quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-139-138.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 kubelet/ip-10-0-139-138.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-router-scheduler-5585754765-zj4ms [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-router-scheduler-5585754765 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-784875bf86 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:48 e2e-test-llm-tls-resources-1d9e6e79 OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy tls-verification-test-kserve-route-authn: failed to get HTTPRoute e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "tls-verification-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Secret e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Deployment e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Service e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.ServiceAccount e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /e2e-test-llm-tls-resources-1d9eb1eecce41c4e293c9684eca957ddbd27 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:57 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Role e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.RoleBinding e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Deployment e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:53:58 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created Created v1.Service e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:26 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:26 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Updated Updated v1.Secret e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:80 2026-07-28 17:54:26 e2e-test-llm-tls-resources-1d9e6e79 LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:172 ### Pod tls-verification-test-kserve-784875bf86-2bg2p (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:212 #### init-container 'storage-initializer' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:228 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '3a537666-5f05-4d2d-8644-414628f8e833', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '261'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.142.94:10250/containerLogs/e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-784875bf86-2bg2p/storage-initializer?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:228 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'd6fcc748-66a0-409b-af2d-d8c1068724d9', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '246'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.142.94:10250/containerLogs/e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-784875bf86-2bg2p/main?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:172 ### Pod tls-verification-test-kserve-router-scheduler-5585754765-zj4ms (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:212 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:228 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '93069984-bfbe-4de4-bfbb-f1107032060b', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Tue, 28 Jul 2026 17:55:07 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.139.138:10250/containerLogs/e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-router-scheduler-5585754765-zj4ms/main?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: eae40c47-3590-46be-bc99-690ffb4b0693 [e2e-llm-inference-service] resourceVersion: '102610' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:54:29Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.133.0.62 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-5585754765-zj4ms [e2e-llm-inference-service] uid: 6414a3ad-a794-486e-aa55-0b735d885216 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 5112918d-a605-490d-94b7-ecdde948ea08 [e2e-llm-inference-service] resourceVersion: '102109' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - notReadyAddresses: [e2e-llm-inference-service] - ip: 10.132.0.50 [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-kserve-784875bf86-2bg2p [e2e-llm-inference-service] uid: 45e91d3b-b60b-4698-95db-f4d181797730 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-784875bf86-2bg2p [e2e-llm-inference-service] generateName: tls-verification-test-kserve-784875bf86- [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 45e91d3b-b60b-4698-95db-f4d181797730 [e2e-llm-inference-service] resourceVersion: '102245' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 784875bf86 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.50/23"],"mac_address":"0a:58:0a:84:00:32","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.50/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.50\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:32\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: tls-verification-test-kserve-784875bf86 [e2e-llm-inference-service] uid: 68dd2c2b-6c38-4bac-befb-4efa4328ef2a [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-142-94 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"68dd2c2b-6c38-4bac-befb-4efa4328ef2a"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TOKIO_WORKER_THREADS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:03Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.50"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-2zt96 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: TOKIO_WORKER_THREADS [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2zt96 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1001150000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to infer\ [e2e-llm-inference-service] \ RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/* 2>/dev/null\n\ [e2e-llm-inference-service] \ grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/* 2>/dev/null\n\ [e2e-llm-inference-service] \n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"$hca_dir\"\ [e2e-llm-inference-service] \ ]; then\n hca_name=$(basename \"$hca_dir\")\n port_state_file=\"\ [e2e-llm-inference-service] $hca_dir/ports/1/state\" # Assume port 1\n type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\ [e2e-llm-inference-service] \n\n echo \"[Infer RoCE] Check if the port state file ${port_state_file}\ [e2e-llm-inference-service] \ exists and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] &&\ [e2e-llm-inference-service] \ grep -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found active\ [e2e-llm-inference-service] \ HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n else\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Skipping inactive or down HCA: $hca_name\"\ [e2e-llm-inference-service] \n fi\n fi\n done\n\n # Check if we found any active HCAs\n\ [e2e-llm-inference-service] \ if [ ${#active_hcas[@]} -gt 0 ]; then\n # Join the array elements with\ [e2e-llm-inference-service] \ a comma\n hca_port_pairs=()\n for hca in \"${active_hcas[@]}\";\ [e2e-llm-inference-service] \ do\n hca_port_pairs+=(\"${hca}:1\")\n done\n\n active_hca_list=$(IFS=,;\ [e2e-llm-inference-service] \ echo \"${active_hcas[*]}\")\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\"\ [e2e-llm-inference-service] )\n echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n\ [e2e-llm-inference-service] \ export UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found. NCCL_IB_HCA\ [e2e-llm-inference-service] \ will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Finding GID_INDEX for each active HCA (SR-IOV compatible)...\"\ [e2e-llm-inference-service] \n\n # For SR-IOV environments, find the most common IPv4 RoCE v2 GID index\ [e2e-llm-inference-service] \ across all HCAs\n declare -A gid_index_count\n declare -A hca_gid_index\n\ [e2e-llm-inference-service] \n for hca_name in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Processing HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for\ [e2e-llm-inference-service] \ this HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"$tpath\"\ [e2e-llm-inference-service] \ 2>/dev/null; then\n idx=$(basename \"$tpath\")\n \ [e2e-llm-inference-service] \ gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n \ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo \"\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Found IPv4 RoCE v2 GID for ${hca_name}:\ [e2e-llm-inference-service] \ index=${idx}, gid=${gid_value}\"\n hca_gid_index[\"${hca_name}\"\ [e2e-llm-inference-service] ]=\"${idx}\"\n gid_index_count[\"${idx}\"]=$((${gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]} + 1))\n break # Use first found IPv4 GID per\ [e2e-llm-inference-service] \ HCA\n fi\n fi\n done\n done\n\n\ [e2e-llm-inference-service] \ # Find the most common GID index (most likely to be consistent across\ [e2e-llm-inference-service] \ nodes)\n best_gid_index=\"\"\n max_count=0\n for idx in \"\ [e2e-llm-inference-service] ${!gid_index_count[@]}\"; do\n count=${gid_index_count[\"${idx}\"]}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n \ [e2e-llm-inference-service] \ if [ $count -gt $max_count ]; then\n max_count=$count\n\ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n #\ [e2e-llm-inference-service] \ Use deterministic fallback if tied - prefer index 3 (SR-IOV standard)\n \ [e2e-llm-inference-service] \ if [ ${#gid_index_count[@]} -gt 1 ]; then\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Multiple GID indices found, selecting most common: ${best_gid_index}\"\n \ [e2e-llm-inference-service] \ # If there's a tie, prefer index 3 as it's most common in SR-IOV setups\n\ [e2e-llm-inference-service] \ if [ -n \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\"\ [e2e-llm-inference-service] \ -eq \"$max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for NCCL,\ [e2e-llm-inference-service] \ NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR: No valid\ [e2e-llm-inference-service] \ IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any HCA.\"\n \ [e2e-llm-inference-service] \ fi\n else\n echo \"[Infer RoCE] No active HCAs found, skipping GID_INDEX\ [e2e-llm-inference-service] \ inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints landed in vLLM\ [e2e-llm-inference-service] \ 0.16.0 (vllm-project/vllm#30011).\n# Older versions still need the blanket\ [e2e-llm-inference-service] \ --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+ ]] &&\ [e2e-llm-inference-service] \ [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort -V | head\ [e2e-llm-inference-service] \ -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout 40\"\ [e2e-llm-inference-service] \nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM 0.22.0+\ [e2e-llm-inference-service] \ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" != *\"\ [e2e-llm-inference-service] --kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"* ]] &&\ [e2e-llm-inference-service] \ [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\"\ [e2e-llm-inference-service] \n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name \"\ [e2e-llm-inference-service] facebook/opt-125m\" \"publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2zt96 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1001150000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c34,c14 [e2e-llm-inference-service] fsGroup: 1001150000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-lsvwr [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:54:02Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] hostIP: 10.0.142.94 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.142.94 [e2e-llm-inference-service] podIP: 10.132.0.50 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.50 [e2e-llm-inference-service] startTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] finishedAt: '2026-07-28T17:54:02Z' [e2e-llm-inference-service] containerID: cri-o://836f778821bdcb013729c8233a3ed83746eefb7966a5545bbd480f513a49b6b6 [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] containerID: cri-o://836f778821bdcb013729c8233a3ed83746eefb7966a5545bbd480f513a49b6b6 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2zt96 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1001150000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1001150000 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T17:54:02Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: docker.io/vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] imageID: docker.io/vllm/vllm-openai-cpu@sha256:a8257c201fd2f696c146615987be61579aa2052367b73c27bd3fb1ab8e90bd4b [e2e-llm-inference-service] containerID: cri-o://de7ca7e8a2ea4723cd518850e88e318b4a304691058d8ad9f8c05de55c5d99e9 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-2zt96 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1001150000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1001150000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-5585754765-zj4ms [e2e-llm-inference-service] generateName: tls-verification-test-kserve-router-scheduler-5585754765- [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 6414a3ad-a794-486e-aa55-0b735d885216 [e2e-llm-inference-service] resourceVersion: '102608' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: '5585754765' [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.62/23"],"mac_address":"0a:58:0a:85:00:3e","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.62/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.62\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:3e\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-5585754765 [e2e-llm-inference-service] uid: 2c1be496-67e3-49cd-934c-853bb95e9d5b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-139-138 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"2c1be496-67e3-49cd-934c-853bb95e9d5b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:29Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.62"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-fjhc9 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type:\ [e2e-llm-inference-service] \ single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n -\ [e2e-llm-inference-service] \ pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: no-hit-lru-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-fjhc9 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1001150000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c34,c14 [e2e-llm-inference-service] fsGroup: 1001150000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-lsvwr [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-x8rqm [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:54:29Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:54:29Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] hostIP: 10.0.139.138 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.139.138 [e2e-llm-inference-service] podIP: 10.133.0.62 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.62 [e2e-llm-inference-service] startTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-llm-d-router-endpoint-picker@sha256:7fd1d61cb5505e026eebab0b9653758779c3dee241e37bea92a0bab961a48c50 [e2e-llm-inference-service] containerID: cri-o://3736fb5425fad97dc099aac5e672cf2d1efccc27ab788bfb57b2201900df5151 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-fjhc9 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1001150000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1001150000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 83652cdd-335b-4887-8b3a-8b06b70bdb61 [e2e-llm-inference-service] resourceVersion: '101989' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: tls-verification-test-epp-sa-dockercfg-x8rqm [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"tls-verification-test-epp-sa-dockercfg-x8rqm"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-lsvwr"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: default-dockercfg-lsvwr [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-x8rqm [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-lsvwr [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-x8rqm [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: bf082d13-ddf9-4ccd-a324-2efcccbe202c [e2e-llm-inference-service] resourceVersion: '102020' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.245.93 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.245.93 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 838c9caf-2114-45fd-87e3-a336131f4ea7 [e2e-llm-inference-service] resourceVersion: '101981' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.182.180 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.182.180 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 2fe4f60d-044f-49a4-896d-27081125c056 [e2e-llm-inference-service] resourceVersion: '102003' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:unavailableReplicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TOKIO_WORKER_THREADS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: TOKIO_WORKER_THREADS [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] unavailableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] reason: ReplicaSetUpdated [e2e-llm-inference-service] message: ReplicaSet "tls-verification-test-kserve-784875bf86" is progressing. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: c586fa94-8dff-4780-a5d2-baf803a053bf [e2e-llm-inference-service] resourceVersion: '102613' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n weight: 3\n - pluginRef: no-hit-lru-scorer\n\ [e2e-llm-inference-service] \ weight: 2\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "tls-verification-test-kserve-router-scheduler-5585754765" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-784875bf86 [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 68dd2c2b-6c38-4bac-befb-4efa4328ef2a [e2e-llm-inference-service] resourceVersion: '102001' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 784875bf86 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: tls-verification-test-kserve [e2e-llm-inference-service] uid: 2fe4f60d-044f-49a4-896d-27081125c056 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"2fe4f60d-044f-49a4-896d-27081125c056"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TOKIO_WORKER_THREADS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 784875bf86 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 784875bf86 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:0a55b7ab9c6c77d762de859d20496ef5a9b41cb492d5728752f25851b1e70916 [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: TOKIO_WORKER_THREADS [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: vllm/vllm-openai-cpu:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\n# --kv-transfer-config with OffloadingConnector requires vLLM\ [e2e-llm-inference-service] \ 0.22.0+ (vllm-project/vllm#40020).\nKV_TRANSFER_ARGS=\"\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.22.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.22.0\" ]; then\n if [[ \"${VLLM_ADDITIONAL_ARGS:-}\"\ [e2e-llm-inference-service] \ != *\"--kv-transfer-config\"* ]] && [[ \"${VLLM_ADDITIONAL_ARGS:-}\" !=\ [e2e-llm-inference-service] \ *\"--kv_transfer_config\"* ]] && [[ \"$*\" != *\"--kv-transfer-config\"\ [e2e-llm-inference-service] * ]] && [[ \"$*\" != *\"--kv_transfer_config\"* ]]; then\n KV_TRANSFER_ARGS=\"\ [e2e-llm-inference-service] \"\n fi\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n ${KV_TRANSFER_ARGS} \\\n \\\n --enable-ssl-refresh \\\n --ssl-certfile\ [e2e-llm-inference-service] \ /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: false [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-5585754765 [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 2c1be496-67e3-49cd-934c-853bb95e9d5b [e2e-llm-inference-service] resourceVersion: '102612' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: '5585754765' [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] uid: c586fa94-8dff-4780-a5d2-baf803a053bf [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"c586fa94-8dff-4780-a5d2-baf803a053bf"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: '5585754765' [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: '5585754765' [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.10.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: quay.io/opendatahub/odh-llm-d-router-endpoint-picker:v0.9.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: llm-d.ai/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n-\ [e2e-llm-inference-service] \ type: single-profile-handler\n- type: queue-scorer\n- type: kv-cache-utilization-scorer\n\ [e2e-llm-inference-service] - type: prefix-cache-scorer\n- type: no-hit-lru-scorer\n- type: max-score-picker\n\ [e2e-llm-inference-service] - parameters:\n scheme: https\n type: metrics-data-source\nschedulingProfiles:\n\ [e2e-llm-inference-service] - name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: kv-cache-utilization-scorer\n weight: 2\n - pluginRef:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n weight: 3\n - pluginRef: no-hit-lru-scorer\n\ [e2e-llm-inference-service] \ weight: 2\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '6' [e2e-llm-inference-service] memory: 16Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 8b33e53c-2289-474d-b4df-1d7244940968 [e2e-llm-inference-service] resourceVersion: '102011' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: e3cbbd8c-6ce1-47b3-a584-8654b389f4e8 [e2e-llm-inference-service] resourceVersion: '102007' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service-q2rdf [e2e-llm-inference-service] generateName: tls-verification-test-epp-service- [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 1c3157d9-6100-4d76-a2b4-40b04ba19123 [e2e-llm-inference-service] resourceVersion: '102611' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: tls-verification-test-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-28T17:54:29Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] uid: bf082d13-ddf9-4ccd-a324-2efcccbe202c [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:54:30Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"bf082d13-ddf9-4ccd-a324-2efcccbe202c"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.133.0.62 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-5585754765-zj4ms [e2e-llm-inference-service] uid: 6414a3ad-a794-486e-aa55-0b735d885216 [e2e-llm-inference-service] nodeName: ip-10-0-139-138.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc-k6x2f [e2e-llm-inference-service] generateName: tls-verification-test-kserve-workload-svc- [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: ce6c0423-f68c-42d6-a4f8-1676b98d4b32 [e2e-llm-inference-service] resourceVersion: '102110' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] uid: 838c9caf-2114-45fd-87e3-a336131f4ea7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"838c9caf-2114-45fd-87e3-a336131f4ea7"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.50 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] serving: false [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-kserve-784875bf86-2bg2p [e2e-llm-inference-service] uid: 45e91d3b-b60b-4698-95db-f4d181797730 [e2e-llm-inference-service] nodeName: ip-10-0-142-94.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-rb [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: 8b33e53c-2289-474d-b4df-1d7244940968 [e2e-llm-inference-service] resourceVersion: '102011' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:e2e-test-llm-tls-resources-1d9e6e79:tls-verification-test-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] uid: e3cbbd8c-6ce1-47b3-a584-8654b389f4e8 [e2e-llm-inference-service] resourceVersion: '102007' [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-28T17:53:57Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - llm-d.ai [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102579' [e2e-llm-inference-service] uid: d007e732-558e-4da1-8841-9c9ba89e66cd [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102579' [e2e-llm-inference-service] uid: d007e732-558e-4da1-8841-9c9ba89e66cd [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/e2e-test-llm-tls-resources-1d9e6e79/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [e2e-test-llm-tls-resources-1d9e6e79/tls-verification-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102557' [e2e-llm-inference-service] uid: 66f68e71-6aeb-45ac-8c67-bc764c78b757 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:59Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-28T17:54:02Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route-authn [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102231' [e2e-llm-inference-service] uid: d5777d14-00b7-4763-b3ba-d1624d1ed9e5 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:00Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-28T17:54:02Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102051' [e2e-llm-inference-service] uid: 4bbe1d3c-53e9-44d8-ab54-311282e0af66 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102564' [e2e-llm-inference-service] uid: 5eb462cc-e779-4a4a-a46e-57b209bf25db [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102059' [e2e-llm-inference-service] uid: 0a0d37df-1bf3-4b35-bf8a-a1c10b58660d [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102051' [e2e-llm-inference-service] uid: 4bbe1d3c-53e9-44d8-ab54-311282e0af66 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102564' [e2e-llm-inference-service] uid: 5eb462cc-e779-4a4a-a46e-57b209bf25db [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102059' [e2e-llm-inference-service] uid: 0a0d37df-1bf3-4b35-bf8a-a1c10b58660d [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102051' [e2e-llm-inference-service] uid: 4bbe1d3c-53e9-44d8-ab54-311282e0af66 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:54:26Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102564' [e2e-llm-inference-service] uid: 5eb462cc-e779-4a4a-a46e-57b209bf25db [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102059' [e2e-llm-inference-service] uid: 0a0d37df-1bf3-4b35-bf8a-a1c10b58660d [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.e2e-test-llm-tls-resources-1d9e6e79.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:264 --- [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:diagnostic.py:265 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"e307c73d-192a-49c2-be59-815fcb232add"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-28T17:53:58Z' [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] namespace: e2e-test-llm-tls-resources-1d9e6e79 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: e307c73d-192a-49c2-be59-815fcb232add [e2e-llm-inference-service] resourceVersion: '102032' [e2e-llm-inference-service] uid: dbcf48b8-a6f0-4588-ae40-b99a8ad34814 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-28T17:55:08.428300] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'e2e-test-llm-tls-resources-1d9e6e79', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [delete_llmisvc] [2026-07-28T17:55:08.443693] end - ❌ 0.015s: ❌ Exception when calling CustomObjectsApi->delete_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '6f15a650-68f2-4230-b410-b714934e4803', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:08 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] WARNING e2e.llmisvc.test_llm_tls:test_llm_tls.py:158 Warning: Failed to cleanup service tls-verification-test: ❌ Exception when calling CustomObjectsApi->delete_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '6f15a650-68f2-4230-b410-b714934e4803', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:08 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_tls_resources] [2026-07-28T17:55:08.443865] end - ❌ 129.056s: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '55f53948-bd64-47b0-a3f6-9ba965c020c6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': '856a9ffa-bb2f-4725-b608-34693201283a', 'X-Kubernetes-Pf-Prioritylevel-Uid': 'd2eb8537-5463-459a-8ae9-7aa31e1ed31f', 'Date': 'Tue, 28 Jul 2026 17:55:06 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.conftest:conftest.py:168 Skipping deletion of namespace e2e-test-llm-tls-resources-1d9e6e79 (SKIP_DELETION_ON_FAILURE) [e2e-llm-inference-service] =============================== warnings summary =============================== [e2e-llm-inference-service] llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-utilization-detector] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] /workspace/source/python/kserve/.venv/lib64/python3.11/site-packages/pytest_asyncio/plugin.py:761: DeprecationWarning: The event_loop fixture provided by pytest-asyncio has been redefined in [e2e-llm-inference-service] /workspace/source/test/e2e/conftest.py:43 [e2e-llm-inference-service] Replacing the event_loop fixture with a custom implementation is deprecated [e2e-llm-inference-service] and will lead to errors in the future. [e2e-llm-inference-service] If you want to request an asyncio event loop with a scope other than function [e2e-llm-inference-service] scope, use the "scope" argument to the asyncio mark when marking the tests. [e2e-llm-inference-service] If you want to return different types of event loops, use the event_loop_policy [e2e-llm-inference-service] fixture. [e2e-llm-inference-service] [e2e-llm-inference-service] warnings.warn( [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-utilization-detector] [e2e-llm-inference-service] llmisvc/test_flow_control.py:47: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_flow_control.py::test_flow_control_smoke[cluster_cpu-cluster_single_node-flow-control-concurrency-detector] [e2e-llm-inference-service] llmisvc/test_flow_control.py:47: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-pd-config-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-v06-nonzero-threshold-migration-workload-llmd-simulator-pd] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-tokenizer-kvcache-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:243: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py:40: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_tls.py:92: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html [e2e-llm-inference-service] - generated xml file: /workspace/artifacts-dir/junit_e2e-llminferenceservice_and_cluster_cpu_and_not_autoscaling_and_not_tracing.xml - [e2e-llm-inference-service] --------------------------------- JSON report ---------------------------------- [e2e-llm-inference-service] report saved to: /workspace/artifacts-dir/e2e_results-llminferenceservice_and_cluster_cpu_and_not_autoscaling_and_not_tracing.json [e2e-llm-inference-service] =========================== short test summary info ============================ [e2e-llm-inference-service] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_deployment[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_hpa_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_keda_lws[cluster_cpu-cluster_multi_node-router-managed-workload-llmd-simulator-lws-prometheus-scrape-scaling-keda] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] FAILED llmisvc/test_llm_autoscaling_wva.py::test_llm_autoscaling_cleanup_hpa[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-no-replicas-prometheus-scrape-scaling-hpa] [e2e-llm-inference-service] FAILED llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] ERROR llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] !!!!!!!!!!!!!!!!!!!!!!!!!! stopping after 10 failures !!!!!!!!!!!!!!!!!!!!!!!!!! [e2e-llm-inference-service] !!!!!!!!!!!! xdist.dsession.Interrupted: stopping after 5 failures !!!!!!!!!!!!! [e2e-llm-inference-service] == 9 failed, 37 passed, 3 skipped, 25 warnings, 1 error in 5612.79s (1:33:32) == [must-gather] [must-gather ] OUT 2026-07-28T17:55:10.648721235Z Using must-gather plug-in image: quay.io/modh/must-gather:rhoai-2.24 [must-gather] When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: [must-gather] ClusterID: 54734bda-8a17-4dff-b565-66b6a98adf6c [must-gather] ClientVersion: 4.21.10 [must-gather] ClusterVersion: Stable at "4.21.26" [must-gather] ClusterOperators: [must-gather] clusteroperator/authentication is missing [must-gather] clusteroperator/cloud-credential is missing [must-gather] clusteroperator/cluster-autoscaler is missing [must-gather] clusteroperator/config-operator is missing [must-gather] clusteroperator/etcd is missing [must-gather] clusteroperator/machine-api is missing [must-gather] clusteroperator/machine-approver is missing [must-gather] clusteroperator/machine-config is missing [must-gather] clusteroperator/marketplace is missing [must-gather] [must-gather] [must-gather] [must-gather ] OUT 2026-07-28T17:55:10.710653868Z namespace/openshift-must-gather-v2h8h created [must-gather] [must-gather ] OUT 2026-07-28T17:55:10.722361606Z clusterrolebinding.rbac.authorization.k8s.io/must-gather-5h5r5 created [must-gather] [must-gather ] OUT 2026-07-28T17:55:10.763277445Z pod for plug-in image quay.io/modh/must-gather:rhoai-2.24 created [must-gather] [must-gather-cd8cw] OUT 2026-07-28T17:57:20.763611483Z gather did not start: resource name may not be empty [must-gather] [must-gather ] OUT 2026-07-28T17:57:20.77721265Z namespace/openshift-must-gather-v2h8h deleted [must-gather] [must-gather] [must-gather] Reprinting Cluster State: [must-gather] When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: [must-gather] ClusterID: 54734bda-8a17-4dff-b565-66b6a98adf6c [must-gather] ClientVersion: 4.21.10 [must-gather] ClusterVersion: Stable at "4.21.26" [must-gather] ClusterOperators: [must-gather] clusteroperator/dns is not available (DNS "default" is unavailable.) because DNS default is degraded [must-gather] clusteroperator/image-registry is not available (Available: The deployment does not have available replicas [must-gather] NodeCADaemonAvailable: The daemon set node-ca does not have available replicas [must-gather] ImagePrunerAvailable: Pruner CronJob has been created) because Degraded: The deployment does not have available replicas [must-gather] clusteroperator/ingress is not available (The "default" ingress controller reports Available=False: IngressControllerUnavailable: One or more status conditions indicate unavailable: DeploymentAvailable=False (DeploymentUnavailable: The deployment has Available status condition set to False (reason: MinimumReplicasUnavailable) with message: Deployment does not have minimum availability.)) because The "default" ingress controller reports Degraded=True: DegradedConditions: One or more other status conditions indicate a degraded state: DeploymentAvailable=False (DeploymentUnavailable: The deployment has Available status condition set to False (reason: MinimumReplicasUnavailable) with message: Deployment does not have minimum availability.), DeploymentReplicasMinAvailable=False (DeploymentMinimumReplicasNotMet: 0/1 of replicas are available, max unavailable is 0: Some pods are not scheduled: Pod "router-default-84ccd6467-wfvnh" cannot be scheduled: 0/1 nodes are available: 1 node(s) had untolerated taint(s). no new claims to deallocate, preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling. Make sure you have sufficient worker nodes.). [must-gather] clusteroperator/network is progressing: DaemonSet "/openshift-multus/multus" is not available (awaiting 1 nodes) [must-gather] DaemonSet "/openshift-multus/multus-additional-cni-plugins" is not available (awaiting 1 nodes) [must-gather] DaemonSet "/openshift-multus/network-metrics-daemon" is not available (awaiting 1 nodes) [must-gather] DaemonSet "/openshift-ovn-kubernetes/ovnkube-node" is not available (awaiting 1 nodes) [must-gather] DaemonSet "/openshift-network-operator/iptables-alerter" is not available (awaiting 1 nodes) [must-gather] Deployment "/openshift-network-console/networking-console-plugin" is not available (awaiting 1 nodes) [must-gather] clusteroperator/node-tuning is not available (DaemonSet "tuned" has no available Pod(s)) because DaemonSet "tuned" available [must-gather] clusteroperator/storage is not available (AWSEBSCSIDriverOperatorCRAvailable: AWSEBSDriverNodeServiceControllerAvailable: Waiting for the DaemonSet to deploy the CSI Node Service) because AWSEBSCSIDriverOperatorCRDegraded: All is well [must-gather] clusteroperator/authentication is missing [must-gather] clusteroperator/cloud-credential is missing [must-gather] clusteroperator/cluster-autoscaler is missing [must-gather] clusteroperator/config-operator is missing [must-gather] clusteroperator/etcd is missing [must-gather] clusteroperator/machine-api is missing [must-gather] clusteroperator/machine-approver is missing [must-gather] clusteroperator/machine-config is missing [must-gather] clusteroperator/marketplace is missing [must-gather] [must-gather] [must-gather] error: gather did not start for pod must-gather-cd8cw: resource name may not be empty [git-push-artifacts] WORK_DIR: /workspace/odh-ci-artifacts [git-push-artifacts] REPO_PATH: opendatahub-io/odh-build-metadata [git-push-artifacts] REPO_BRANCH: ci-artifacts [git-push-artifacts] SPARSE_FILE_PATH: test-artifacts/docs [git-push-artifacts] SOURCE_PATH: /workspace/artifacts-dir [git-push-artifacts] DEST_PATH: test-artifacts/kserve-group-test-hbdjx [git-push-artifacts] ALWAYS_PASS: false [git-push-artifacts] configuring gh token [git-push-artifacts] taking github token from Konflux bot [git-push-artifacts] Initialized empty Git repository in /workspace/odh-ci-artifacts/.git/ [git-push-artifacts] Using partial fetch with sparse checkout for: test-artifacts/docs [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] * [new branch] ci-artifacts -> origin/ci-artifacts [git-push-artifacts] Already on 'ci-artifacts' [git-push-artifacts] branch 'ci-artifacts' set up to track 'origin/ci-artifacts'. [git-push-artifacts] TASK_NAME=kserve-group-test-hbdjx-e2e-llm-inference-service [git-push-artifacts] PIPELINERUN_NAME=kserve-group-test-hbdjx [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] Already up to date. [git-push-artifacts] -rw-r--r--. 1 root 1001540000 84559 Jul 28 17:57 /workspace/odh-ci-artifacts/test-artifacts/kserve-group-test-hbdjx/e2e-llm-inference-service.tar.gz [git-push-artifacts] [ci-artifacts be6b496] Updating CI Artifacts in e2e-llm-inference-service [git-push-artifacts] 1 file changed, 0 insertions(+), 0 deletions(-) [git-push-artifacts] create mode 100644 test-artifacts/kserve-group-test-hbdjx/e2e-llm-inference-service.tar.gz [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] Already up to date. [git-push-artifacts] To https://github.com/opendatahub-io/odh-build-metadata.git [git-push-artifacts] acd34d9..be6b496 ci-artifacts -> ci-artifacts [fail-if-needed] Failing pipeline because deploy-and-e2e step failed container step-fail-if-needed has failed : [{"key":"StartedAt","value":"2026-07-28T17:57:25.830Z","type":3}]